site reliability engineer

site reliability engineer

Full-Time 60000 - 80000 £ / year (est.) No working from home possible
Enfint

At a Glance

  • Tasks: Lead a global SRE team, ensuring system reliability and performance through innovative practices.
  • Company: EPAM, a leader in digital engineering and AI transformation services.
  • Benefits: Competitive salary, health benefits, remote work options, and opportunities for professional growth.
  • Other info: Work in a fast-paced environment with excellent career advancement opportunities.
  • Why this job: Join a dynamic team to drive innovation and make a real impact in tech.
  • Qualifications: Experience in Site Reliability Engineering or DevOps, with strong automation skills.

The predicted salary is between 60000 - 80000 £ per year.

EPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.

Responsibilities:

  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment.
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices.
  • Define and monitor KPIs for system reliability, performance, and operational efficiency.
  • Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques.
  • Develop robust incident management frameworks and lead major incident response activities for critical systems.
  • Implement blameless postmortems and deliver systemic improvements across production environments.
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems.
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services.
  • Drive resilience strategies with highly available architectures and disaster recovery readiness.
  • Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes.

Requirements:

  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments.
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights.
  • Hands-on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices.
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles.
  • Knowledge of resilience design patterns, high availability, and fault-tolerant architectures.
  • Familiarity with AI/ML-driven approaches for operational efficiency and system reliability.
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture.

Nice to have:

  • Experience in financial services or other highly regulated, mission-critical environments.
  • Certifications in cloud technologies such as AWS.
  • Exposure to AIOps platforms or advanced observability tooling.

site reliability engineer employer: Enfint

EPAM is an exceptional employer for Site Reliability Engineers, offering a dynamic work culture that prioritises engineering excellence and team empowerment. With a strong focus on employee growth through collaboration across diverse teams and the adoption of cutting-edge technologies, EPAM provides unique opportunities to lead transformative projects in a global setting, all while fostering an automation-first mindset that enhances operational efficiency.

Enfint

Contact Details:

Enfint Recruitment Team

We think you need these skills to ace site reliability engineer

Site Reliability Engineering
DevOps
Platform Operations
Observability Platforms
Troubleshooting Distributed Systems
Telemetry-Driven Insights
Automation