Back to SearchLooking for something else? Find a vacancy that works for you.
Considering applying for this job Do not delay, scroll down and make your application as soon as possible to avoid missing out.
Send us your CV to receive a personalized offer. Find me a jobWe're looking for a Senior Site Reliability Engineer β Observability to join our team in London, United Kingdom in a hybrid working mode.
You will be part of the Production Engineering β Observability team, driving the strategic initiative to implement and expand a modern observability platform built on OpenTelemetry and OpenSearch.
This program focuses on enhancing monitoring, resilience and operational stability for critical FIC trading systems by enabling faster incident detection, reducing outage duration and providing actionable insights across the technology estate.
This is a hands-on engineering role where you'll define observability standards and deliver enterprise-scale solutions while promoting best practices for operational excellence.
Gather requirements and perform analysis of existing monitoring and observability platformsDefine observability standards, telemetry strategies and alerting frameworksImplement OpenTelemetry-based instrumentation and OpenSearch solutions across applications and infrastructureDesign dashboards, analytics and reporting to improve transparency and operational efficiencyDevelop automation tools and processes to enhance observability and reduce manual overheadIntegrate observability frameworks with enterprise monitoring platforms such as GeneosProvide documentation and operational handover to ensure long-term sustainabilityApply Site Reliability Engineering principles to drive stability, scalability and incident reductionProven experience as Senior SRE or Observability Engineer implementing enterprise-scale observability solutionsStrong expertise in OpenTelemetry including instrumentation and telemetry xohmjla pipelinesIn-depth knowledge of OpenSearch for architecture, data indexing, optimisation and analyticsExperience developing dashboards and alerts using GrafanaFamiliarity with enterprise monitoring tools such as Geneos and related observability technologiesPractical understanding of Site Reliability Engineering practices and automation approachesBackground in high-availability or mission-critical environments; financial services experience is highly desirableKnowledge of anomaly detection, alert correlation, and incident response automationExposure to observability in cloud-native or hybrid architectures
Senior Site Reliability Engineer - OpenTelemetry in Vauxhall employer: EPAM Systems
Join a dynamic and innovative team in London where your expertise as a Delivery Director will be valued and nurtured. We pride ourselves on a collaborative work culture that fosters continuous improvement and professional growth, offering you the chance to work alongside award-winning teams in the Retail/Consumer Products industry. With a focus on client satisfaction and strategic project delivery, you'll find meaningful and rewarding employment in an environment that champions creativity and technical excellence.