Senior Site Reliability Engineer - OpenTelemetry in London

Senior Site Reliability Engineer - OpenTelemetry in London

London Full-Time On-site
E

We're looking for a Senior Site Reliability Engineer – Observability to join our team in London, United Kingdom in a hybrid working mode. You will be part of the Production Engineering – Observability team, driving the strategic initiative to implement and expand a modern observability platform built on OpenTelemetry and OpenSearch. This program focuses on enhancing monitoring, resilience and operational stability for critical FIC trading systems by enabling faster incident detection, reducing outage duration and providing actionable insights across the technology estate. This is a hands-on engineering role where you’ll define observability standards and deliver enterprise-scale solutions while promoting best practices for operational excellence.ResponsibilitiesGather requirements and perform analysis of existing monitoring and observability platformsDefine observability standards, telemetry strategies and alerting frameworksImplement OpenTelemetry-based instrumentation and OpenSearch solutions across applications and infrastructureDesign dashboards, analytics and reporting to improve transparency and operational efficiencyDevelop automation tools and processes to enhance observability and reduce manual overheadIntegrate observability frameworks with enterprise monitoring platforms such as GeneosProvide documentation and operational handover to ensure long-term sustainabilityApply Site Reliability Engineering principles to drive stability, scalability and incident reductionRequirementsProven experience as Senior SRE or Observability Engineer implementing enterprise-scale observability solutionsStrong expertise in OpenTelemetry including instrumentation and telemetry pipelinesIn-depth knowledge of OpenSearch for architecture, data indexing, optimisation and analyticsExperience developing dashboards and alerts using GrafanaFamiliarity with enterprise monitoring tools such as Geneos and related observability technologiesPractical understanding of Site Reliability Engineering practices and automation approachesBackground in high-availability or mission-critical environments; financial services experience is highly desirableNice to haveKnowledge of anomaly detection, alert correlation, and incident response automationExposure to observability in cloud-native or hybrid architecturesWe offerEPAM Employee Stock Purchase Plan (ESPP)Protection benefits including life assurance, income protection and critical illness coverPrivate medical insurance and dental careEmployee Assistance ProgramCompetitive group pension planCyclescheme, Techscheme and season ticket loansVarious perks such as free Wednesday lunch in-office, on-site massages and regular social eventsLearning and development opportunities including in-house training and coaching, professional certifications, and coursesIf otherwise eligible, participation in the discretionary annual bonus programIf otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program
#J-18808-Ljbffr

Senior Site Reliability Engineer - OpenTelemetry in London employer: EPAM Systems, Inc.

EPAM Systems, Inc. is an exceptional employer that fosters a collaborative and innovative work culture in the heart of London. With a strong focus on employee growth, you will have ample opportunities to enhance your skills through mentorship and cutting-edge projects in cloud-native data solutions. The company also prioritises work-life balance and offers competitive benefits, making it an ideal place for professionals seeking meaningful and rewarding careers in technology.

E

Contact Details:

EPAM Systems, Inc. Recruitment Team