Director of Site Reliability Engineering in London

Director of Site Reliability Engineering in London

London Full-Time 60000 - 80000 £ / year (est.) No working from home possible
EPAM Systems

At a Glance

  • Tasks: Lead a global SRE team, driving reliability and operational excellence with AI solutions.
  • Company: Join a forward-thinking tech company in London with a hybrid work culture.
  • Benefits: Enjoy competitive pay, health coverage, stock options, and fun perks like free lunches.
  • Other info: Great opportunities for learning and career growth in a dynamic environment.
  • Why this job: Make a real impact on critical systems while shaping the future of technology.
  • Qualifications: Strong background in SRE, DevOps, and automation; leadership skills are a must.

The predicted salary is between 60000 - 80000 £ per year.

We're looking for a Director of Site Reliability Engineering to join our team in London, United Kingdom in a hybrid working mode. This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands‐on governance to ensure highly available, resilient systems that align with business and regulatory requirements. As a technology thought leader, you will influence engineering standards, enhance operational frameworks, and foster a culture of continuous improvement across mission‐critical environments.

Responsibilities

  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self‐healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation‐first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes

Requirements

  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights
  • Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles
  • Knowledge of resilience design patterns, high availability, and fault‐tolerant architectures
  • Familiarity with AI/ML-driven approaches for operational efficiency and system reliability
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture

Nice to have

  • Experience in financial services or other highly regulated, mission‐critical environments
  • Certifications in cloud technologies, such as AWS
  • Exposure to AIOps platforms or advanced observability tooling

We offer

  • EPAM Employee Stock Purchase Plan (ESPP)
  • Protection benefits including life assurance, income protection and critical illness cover
  • Private medical insurance and dental care
  • Employee Assistance Program
  • Cyclescheme, Techscheme and season ticket loans
  • Various perks such as free Wednesday lunch in-office, on‐site massages and regular social events
  • Learning and development opportunities including in‐house training and coaching, professional certifications, and courses
  • If otherwise eligible, participation in the discretionary annual bonus program
  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long‐Term Incentive (LTI) Program

Director of Site Reliability Engineering in London employer: EPAM Systems

EPAM Systems is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the rapidly evolving field of AI. With a strong focus on employee growth, you will have access to extensive learning opportunities and benefits such as stock purchase plans, all while working remotely or in a hybrid model across Europe. Join us to make a meaningful impact in the energy and utilities sector, driving transformation strategies alongside industry leaders.

EPAM Systems

Contact Details:

EPAM Systems Recruitment Team

We think you need these skills to ace Director of Site Reliability Engineering in London

Site Reliability Engineering
DevOps
Platform Operations
Observability Platforms
Troubleshooting Distributed Systems
Telemetry-Driven Insights
Automation