Job Description
Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office)
\nRole Overview
\nWe're recruiting for an experienced Observability SME to define, implement, and govern enterprise-wide observability capabilities for cloud-native and distributed applications running on Microsoft Azure, for a leading organisation. The role requires establishing observability standards, ensuring end-to-end visibility across business-critical platforms, and enabling proactive monitoring, faster incident resolution, and improved platform reliability through modern observability practices - with deep expertise in Grafana, OpenTelemetry, distributed tracing, SRE, event-driven architecture, and Azure Integration Services.
\nKey Responsibilities
\n- \n
- Define and implement enterprise observability strategies, standards, and governance frameworks \n
- Design and manage observability solutions covering metrics, logs, traces, and application telemetry \n
- Establish monitoring, alerting, and diagnostics best practices across cloud-native platforms and microservices \n
- Design and implement distributed tracing solutions using OpenTelemetry and modern observability tools \n
- Develop and maintain Grafana dashboards for engineering, operations, business, and leadership stakeholders \n
- Monitor platform performance, availability, reliability, and service health across Azure services \n
- Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs \n
- Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience \n
- Drive root cause analysis, incident investigations, and continuous service improvement initiatives \n
- Champion operational excellence through proactive monitoring, automation, and reliability engineering practices \n
What You Will Ideally Bring
\n- \n
- Strong experience in enterprise observability, monitoring, and operational support \n
- Expertise in Grafana dashboard development and observability platform management \n
- Hands-on experience with OpenTelemetry, distributed tracing, and telemetry frameworks \n
- Strong understanding of Site Reliability Engineering (SRE) principles and practices \n
- Experience monitoring cloud-native applications, microservices, and distributed systems \n
- Knowledge of Azure monitoring services, diagnostics, and observability ecosystems (Azure Monitor, Log Analytics, Event Hub, Service Bus, Functions) \n
- Experience implementing alerting strategies, incident management, and performance monitoring solutions \n
- Strong understanding of Event-Driven Architecture and asynchronous application behaviour \n
- Experience defining and measuring performance, reliability, and availability metrics \n
- Excellent problem-solving, analytical, stakeholder management, and communication skills \n
- Cosmos DB and/or PostgreSQL experience (desirable) \n
Contract Details
\n- \n
- Duration: Initial 12 months \n
- Location: 3 days onsite in London \n
- Rate: £450-475/day (inside) \n
Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office) employer: Hamilton Barnes
Hamilton Barnes is an exceptional employer, offering a dynamic work environment in London where innovation meets opportunity. With competitive salaries, night shift allowances, and a strong focus on employee development through ongoing training, we empower our Field Service Engineers to advance their careers while enjoying a supportive team culture. Join us to be part of a growing engineering team that values your skills and fosters professional growth.