Observability SME in London

Observability SME in London

London Full-Time No working from home possible
Tcs Uk

The Role

We are seeking an experienced Observability SME to define, implement, and govern enterprise-wide observability capabilities for cloud-native and distributed applications running on Microsoft Azure. The ideal candidate will possess deep expertise in Grafana, OpenTelemetry, Monitoring, Alerting, Distributed Tracing, Site Reliability Engineering (SRE), Event-Driven Architecture, Azure Integration Services, and Operational Excellence.

The role requires establishing observability standards, ensuring end-to-end visibility across business-critical platforms, and enabling proactive monitoring, faster incident resolution, and improved platform reliability through modern observability practices.

Your Responsibilities

Define and implement enterprise observability strategies, standards, and governance frameworks.

Design and manage observability solutions covering metrics, logs, traces, and application telemetry.

Establish monitoring, alerting, and diagnostics best practices across cloud-native platforms and microservices.

Design and implement distributed tracing solutions using OpenTelemetry and modern observability tools.

Develop and maintain Grafana dashboards for engineering, operations, business, and leadership stakeholders.

Monitor platform performance, availability, reliability, and service health across Azure services.

Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs.

Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience.

Drive root cause analysis, incident investigations, and continuous service improvement initiatives.

Champion operational excellence through proactive monitoring, automation, and reliability engineering practices.

Your Profile

Essential Skills/Knowledge/Experience

Strong experience in enterprise observability, monitoring, and operational support.

Expertise in Grafana dashboard development and observability platform management.

Hands-on experience with OpenTelemetry, Distributed Tracing, and telemetry frameworks.

Strong understanding of Site Reliability Engineering (SRE) principles and practices.

Experience monitoring cloud-native applications, microservices, and distributed systems.

Knowledge of Azure monitoring services, diagnostics, and observability ecosystems.

Experience implementing alerting strategies, incident management, and performance monitoring solutions.

Strong understanding of Event-Driven Architecture and asynchronous application behaviour.

Experience defining and measuring performance, reliability, and availability metrics.

Excellent problem-solving, analytical, stakeholder management, and communication skills.

Desirable Skills/Knowledge/Experience

Grafana

OpenTelemetry

Azure Monitor & Log Analytics

Azure Event Hub

Azure Service Bus

Azure Functions

Cosmos DB

PostgreSQL

Microservices Architecture

Event-Driven Architecture

Distributed Systems

Application Performance Monitoring (APM)

DevOps and CI/CD practices

Cloud-Native Platform Engineering

Reliability Engineering and Automation

Root Cause Analysis (RCA) and Incident Management

Azure Cloud Services and Integration Platforms

Observability Governance and Operational Excellence

Observability SME in London employer: Tcs Uk

As a Data Governance Specialist at our company, you will be part of a dynamic team dedicated to driving customer excellence through robust data governance practices. We pride ourselves on fostering a collaborative work culture that values innovation and continuous learning, offering ample opportunities for professional growth in a supportive environment. Located in a vibrant area, our workplace not only provides a stimulating atmosphere but also encourages a healthy work-life balance, making it an ideal place for those seeking meaningful and rewarding employment.

Tcs Uk

Contact Details:

Tcs Uk Recruitment Team