The Role
We are seeking an experienced Observability SME to define, implement, and govern enterprise-wide observability capabilities for cloud-native and distributed applications running on Microsoft Azure. The ideal candidate will possess deep expertise in Grafana, OpenTelemetry, Monitoring, Alerting, Distributed Tracing, Site Reliability Engineering (SRE), Event-Driven Architecture, Azure Integration Services, and Operational Excellence.
The role requires establishing observability standards, ensuring end-to-end visibility across business-critical platforms, and enabling proactive monitoring, faster incident resolution, and improved platform reliability through modern observability practices.
Your Responsibilities
Define and implement enterprise observability strategies, standards, and governance frameworks.
Design and manage observability solutions covering metrics, logs, traces, and application telemetry.
Establish monitoring, alerting, and diagnostics best practices across cloud-native platforms and microservices.
Design and implement distributed tracing solutions using OpenTelemetry and modern observability tools.
Develop and maintain Grafana dashboards for engineering, operations, business, and leadership stakeholders.
Monitor platform performance, availability, reliability, and service health across Azure services.
Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs.
Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience.
Drive root cause analysis, incident investigations, and continuous service improvement initiatives.
Champion operational excellence through proactive monitoring, automation, and reliability engineering practices.
Your Profile
Essential Skills/Knowledge/Experience
Strong experience in enterprise observability, monitoring, and operational support.
Expertise in Grafana dashboard development and observability platform management.
Hands-on experience with OpenTelemetry, Distributed Tracing, and telemetry frameworks.
Strong understanding of Site Reliability Engineering (SRE) principles and practices.
Experience monitoring cloud-native applications, microservices, and distributed systems.
Knowledge of Azure monitoring services, diagnostics, and observability ecosystems.
Experience implementing alerting strategies, incident management, and performance monitoring solutions.
Strong understanding of Event-Driven Architecture and asynchronous application behaviour.
Experience defining and measuring performance, reliability, and availability metrics.
Excellent problem-solving, analytical, stakeholder management, and communication skills.
Desirable Skills/Knowledge/Experience
Grafana
OpenTelemetry
Azure Monitor & Log Analytics
Azure Event Hub
Azure Service Bus
Azure Functions
Cosmos DB
PostgreSQL
Microservices Architecture
Event-Driven Architecture
Distributed Systems
Application Performance Monitoring (APM)
DevOps and CI/CD practices
Cloud-Native Platform Engineering
Reliability Engineering and Automation
Root Cause Analysis (RCA) and Incident Management
Azure Cloud Services and Integration Platforms
Observability Governance and Operational Excellence
Observability SME in London employer: Tcs Uk
As a Data Governance Specialist at our company, you will be part of a dynamic team dedicated to driving customer excellence through robust data governance practices. We pride ourselves on fostering a collaborative work culture that values innovation and continuous learning, offering ample opportunities for professional growth in a supportive environment. Located in a vibrant area, our workplace not only provides a stimulating atmosphere but also encourages a healthy work-life balance, making it an ideal place for those seeking meaningful and rewarding employment.