At a Glance
- Tasks: Lead a global SRE team, ensuring system reliability and performance through innovative practices.
- Company: EPAM, a leader in digital engineering and AI transformation services.
- Benefits: Competitive salary, health benefits, remote work options, and opportunities for professional growth.
- Other info: Work in a fast-paced environment with excellent career advancement opportunities.
- Why this job: Join a dynamic team to drive innovation and make a real impact in tech.
- Qualifications: Experience in Site Reliability Engineering or DevOps, with strong automation skills.
The predicted salary is between 60000 - 80000 £ per year.
EPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.
Responsibilities:
- Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment.
- Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices.
- Define and monitor KPIs for system reliability, performance, and operational efficiency.
- Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques.
- Develop robust incident management frameworks and lead major incident response activities for critical systems.
- Implement blameless postmortems and deliver systemic improvements across production environments.
- Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems.
- Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services.
- Drive resilience strategies with highly available architectures and disaster recovery readiness.
- Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes.
Requirements:
- Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments.
- Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights.
- Hands-on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices.
- Deep understanding of incident management processes, ITSM standards, and ITIL principles.
- Knowledge of resilience design patterns, high availability, and fault-tolerant architectures.
- Familiarity with AI/ML-driven approaches for operational efficiency and system reliability.
- Ability to lead transformation, influence across teams, and foster continuous improvement in culture.
Nice to have:
- Experience in financial services or other highly regulated, mission-critical environments.
- Certifications in cloud technologies such as AWS.
- Exposure to AIOps platforms or advanced observability tooling.
site reliability engineer employer: Enfint
EPAM is an exceptional employer for Site Reliability Engineers, offering a dynamic work culture that prioritises engineering excellence and team empowerment. With a strong focus on employee growth through collaboration across diverse teams and the adoption of cutting-edge technologies, EPAM provides unique opportunities to lead transformative projects in a global setting, all while fostering an automation-first mindset that enhances operational efficiency.
We think you need these skills to ace site reliability engineer
Site Reliability Engineering
DevOps
Platform Operations
Observability Platforms
Troubleshooting Distributed Systems
Telemetry-Driven Insights
Automation