Site Reliability Engineer with 5+ years of experience operating and automating mission-critical Linux infrastructure at scale (2,800+ servers) for core banking systems.
- Automation: Ansible-driven infrastructure automation
- Observability: Prometheus, Grafana, ELK Stack
- Incident Management: Structured incident response practices
- Infrastructure: Large-scale Linux systems (enterprise / banking environments)
- CI/CD: Pipeline adoption and operational improvement
- Maintained 99.9% system availability
- Reduced operational toil by ~50%
- Improved MTTR through:
- Proactive monitoring
- Runbook development
- CI/CD adoption
- LPIC-2 Certified
Site Reliability Engineer (SRE), Platform Engineering, and DevOps roles across the EU.