Director of Site Reliability Engineering

Director of Site Reliability Engineering

Full-Time 60000 - 80000 £ / year (est.) Home office (partial)
EPAM Systems, Inc.

At a Glance

  • Tasks: Lead a global team to enhance system reliability and drive automation using AI.
  • Company: Join a forward-thinking tech company in London with a hybrid work model.
  • Benefits: Enjoy competitive salary, health benefits, and perks like free lunches and on-site massages.
  • Other info: Great opportunities for learning, development, and career growth in a dynamic environment.
  • Why this job: Make a real impact by shaping the future of technology and operational excellence.
  • Qualifications: Strong background in Site Reliability Engineering and expertise in automation and incident management.

The predicted salary is between 60000 - 80000 £ per year.

We're looking for a Director of Site Reliability Engineering to join our team in London, United Kingdom in a hybrid working mode. This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands-on governance to ensure highly available, resilient systems that align with business and regulatory requirements. As a technology thought leader, you will influence engineering standards, enhance operational frameworks, and foster a culture of continuous improvement across mission-critical environments.

Responsibilities

  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes

Requirements

  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights
  • Hands-on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles
  • Knowledge of resilience design patterns, high availability, and fault-tolerant architectures
  • Familiarity with AI/ML-driven approaches for operational efficiency and system reliability
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture

Nice to have

  • Experience in financial services or other highly regulated, mission-critical environments
  • Certifications in cloud technologies, such as AWS
  • Exposure to AIOps platforms or advanced observability tooling

We offer

  • EPAM Employee Stock Purchase Plan (ESPP)
  • Protection benefits including life assurance, income protection and critical illness cover
  • Private medical insurance and dental care
  • Employee Assistance Program
  • Competitive group pension plan
  • Cyclescheme, Techscheme and season ticket loans
  • Various perks such as free Wednesday lunch in-office, on-site massages and regular social events
  • Learning and development opportunities including in-house training and coaching, professional certifications, and courses
  • If otherwise eligible, participation in the discretionary annual bonus program
  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program

Director of Site Reliability Engineering employer: EPAM Systems, Inc.

EPAM Systems, Inc. is an exceptional employer that fosters a dynamic and inclusive work culture, offering employees the flexibility of a hybrid working model in the vibrant city of London. With a strong focus on professional development, team members are encouraged to grow their skills and advance their careers while contributing to impactful digital transformation projects in the energy sector. Joining EPAM means being part of a forward-thinking company that values innovation and collaboration, making it an ideal place for those seeking meaningful and rewarding employment.

EPAM Systems, Inc.

Contact Details:

EPAM Systems, Inc. Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Director of Site Reliability Engineering

Join the IT Consultancy Buzz

Get involved in local or virtual IT consultancy meetups and forums. This is where we can rub shoulders with industry professionals, get insights into what EPAM Systems, Inc. values, and even spot unadvertised opportunities. Don't miss out on these chances to make a name for ourselves in the IT world!

Show Off Your Skills

Create a personal project or case study relevant to the challenges EPAM Systems, Inc. might face. Use platforms like GitHub or Medium to share your findings. This not only demonstrates our consulting skills but shows a proactive attitude, making us stand out from the crowd when applying for that full-time gig.

Leverage LinkedIn for Connections

Follow and engage with the relevant thought leaders and influencers in IT consultancy on LinkedIn. Share insightful content and join discussions to gain visibility. A well-placed comment or shared article could catch the attention of someone at EPAM Systems, Inc.!

Direct Apply to EPAM Systems, Inc.

Let's not forget to apply directly through the EPAM Systems, Inc. website! Tailor your application to showcase our understanding of their consulting style and how we can contribute to their projects. A personalised approach can make a huge difference in landing that full-time position!

We think you need these skills to ace Director of Site Reliability Engineering

Site Reliability Engineering
DevOps
Platform Operations
Observability Platforms
Distributed Systems Troubleshooting
Telemetry-Driven Insights
Automation

Some tips for your application 🫡

Showcase Your Problem-Solving Skills:In IT consulting, it's all about problem-solving, so make sure your CV highlights your analytical skills and any relevant projects you've tackled. Mention specific technologies or methodologies you've used to resolve issues or improve processes; this shows you can think critically and deliver results, which is vital for us at EPAM Systems, Inc..

Highlight Relevant Certifications:Certifications like ITIL, PMP, or even specific tech stack qualifications can really make you stand out. Make sure to include these in your CV, as they not only demonstrate your expertise but also your commitment to staying current in the field. We love seeing candidates who are proactive about their professional development!

Tailor Your Cover Letter:Your cover letter is your chance to connect personally with us at EPAM Systems, Inc.. Share stories about your experiences in IT consulting, and how they shaped your desire to join our team. Mention why you’re excited about this particular role, and how you see yourself contributing to our projects.

Keep It Clear and Concise:We're all busy, so make sure your application is easy to read. Use bullet points for key achievements, and don’t overload us with jargon. A clean, professional layout goes a long way. Remember, the clearer your application, the more likely we are to invite you in for an interview!

How to prepare for a job interview at EPAM Systems, Inc.

Brush Up on Your Technical Skills

For an IT consulting role, be ready to demonstrate your technical prowess. You might face questions on systems integration, cloud technologies, or even troubleshooting specific software. If you have experience with tools like AWS, Azure, or even specific programming languages, make sure you can talk about them fluently.

Showcase Your Problem-Solving Approach

IT consulting is all about solving problems for clients. Think about how you can illustrate your approach to a past challenge using the STAR method (Situation, Task, Action, Result). It's a great way to show how you tackle complex issues and come up with effective solutions.

Know the Business Impact of IT Solutions

When discussing your experiences, focus not just on the tech solutions you implemented, but also on their business impact. Employers want to see that you can connect IT with organisational goals. Prep examples that highlight how your tech contributions improved efficiency or reduced costs for past clients or projects.

Prepare for Behavioural Questions

Since IT consulting often involves teamwork and client interactions, expect behavioural questions that assess your interpersonal skills. Be prepared with examples that demonstrate your adaptability, communication skills, and how you handle client feedback. Before the interview, think of situations where you worked closely with clients to create effective IT strategies or changes.