At a Glance
- Tasks: Lead a global SRE team, driving reliability and operational excellence with AI solutions.
- Company: Join a forward-thinking tech company in London with a hybrid work culture.
- Benefits: Enjoy competitive pay, health coverage, stock options, and fun perks like free lunches.
- Other info: Great opportunities for learning and career growth in a dynamic environment.
- Why this job: Make a real impact on critical systems while shaping the future of technology.
- Qualifications: Strong background in SRE, DevOps, and automation; leadership skills are a must.
The predicted salary is between 60000 - 80000 £ per year.
We're looking for a Director of Site Reliability Engineering to join our team in London, United Kingdom in a hybrid working mode. This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands‐on governance to ensure highly available, resilient systems that align with business and regulatory requirements. As a technology thought leader, you will influence engineering standards, enhance operational frameworks, and foster a culture of continuous improvement across mission‐critical environments.
Responsibilities
- Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
- Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
- Define and monitor KPIs for system reliability, performance, and operational efficiency
- Advance automation, Infrastructure as Code approaches, and promote self‐healing systems using AI/ML techniques
- Develop robust incident management frameworks and lead major incident response activities for critical systems
- Implement blameless postmortems and deliver systemic improvements across production environments
- Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
- Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
- Drive resilience strategies with highly available architectures and disaster recovery readiness
- Champion an automation‐first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes
Requirements
- Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments
- Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights
- Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices
- Deep understanding of incident management processes, ITSM standards, and ITIL principles
- Knowledge of resilience design patterns, high availability, and fault‐tolerant architectures
- Familiarity with AI/ML-driven approaches for operational efficiency and system reliability
- Ability to lead transformation, influence across teams, and foster continuous improvement in culture
Nice to have
- Experience in financial services or other highly regulated, mission‐critical environments
- Certifications in cloud technologies, such as AWS
- Exposure to AIOps platforms or advanced observability tooling
We offer
- EPAM Employee Stock Purchase Plan (ESPP)
- Protection benefits including life assurance, income protection and critical illness cover
- Private medical insurance and dental care
- Employee Assistance Program
- Cyclescheme, Techscheme and season ticket loans
- Various perks such as free Wednesday lunch in-office, on‐site massages and regular social events
- Learning and development opportunities including in‐house training and coaching, professional certifications, and courses
- If otherwise eligible, participation in the discretionary annual bonus program
- If otherwise eligible and hired into a qualifying level, participation in the discretionary Long‐Term Incentive (LTI) Program
Director of Site Reliability Engineering in London employer: EPAM Systems
EPAM Systems is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the rapidly evolving field of AI. With a strong focus on employee growth, you will have access to extensive learning opportunities and benefits such as stock purchase plans, all while working remotely or in a hybrid model across Europe. Join us to make a meaningful impact in the energy and utilities sector, driving transformation strategies alongside industry leaders.