Site Reliability Engineering (SRE) / Observability Technical Lead

Site Reliability Engineering (SRE) / Observability Technical Lead

Full-Time 60000 - 80000 £ / year (est.) No working from home possible
NTT DATA

At a Glance

  • Tasks: Lead observability and reliability projects, ensuring system performance and scalability.
  • Company: Global tech company focused on innovation and collaboration.
  • Benefits: Flexible work options, tailored benefits, and continuous learning opportunities.
  • Other info: Join a diverse team committed to mutual respect and continuous growth.
  • Why this job: Make a real impact in a dynamic environment while developing your leadership skills.
  • Qualifications: 5+ years in SRE or DevOps with expertise in APM and IaC.

The predicted salary is between 60000 - 80000 £ per year.

The team you'; ll be working with: We are seeking an experienced Site Reliability Engineer (SRE) / Observability Technical Lead to join our team and drive the strategy and execution of observability and reliability projects across our clients.

The ideal candidate will have deep expertise in Application Performance Monitoring (APM), Infrastructure as Code (Ia C), automation, and distributed tracing using Open Telemetry.

As a lead, you will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and Dev Ops.

What you'; ll be doing: Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business goals and technical requirements.

Design and implementation of monitoring and observability solutions, collaborating with engineering teams to define standards and best practices.

Manage Infrastructure as Code (Ia C) initiatives using Terraform, coordinating with cloud and infrastructure teams to ensure scalable and secure deployments.

Drive automation strategies for monitoring, alerting, and logging pipelines, focusing on process improvements and operational efficiency.

Develop and maintain comprehensive observability roadmaps, including distributed tracing, logging, and metrics collection strategies.

Collaborate with product management, sales, and pre-sales teams to provide technical expertise and support during solution design and customer engagements.

Lead cross-functional teams to enhance CI/CD pipelines and deployment reliability, ensuring smooth integration of observability tools and practices.

Engage with vendors and strategic partners to evaluate, select, and integrate observability and monitoring solutions, ensuring alignment with organizational needs and fostering strong collaborative relationships.

Mentor and develop junior engineers and analysts, fostering a culture of reliability, observability, and operational excellence.

What experience you'; ll bring: 5+ years of experience in SRE, Observability, or Dev Ops roles, with leadership responsibilities.

Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, App Dynamics, or Dynatrace.

Hands-on experience with Open Telemetry (OTel) for distributed tracing and observability instrumentation.

Strong proficiency in Infrastructure as Code (Ia C) using Terraform.

Solid understanding of cloud platforms including AWS, GCP, or Azure.

Experience with automation/configuration management tools like Ansible, Chef, or Puppet.

Deep knowledge of CI/CD pipelines and tools such as Git Hub Actions, Jenkins, or Azure Dev Ops.

Experience managing Kubernetes and containerized environments (Docker, Helm).

Familiarity with log aggregation and analysis platforms like ELK Stack or Splunk.

Excellent leadership, communication, and collaboration skills.

Who we are: We’re a business with a global reach that empowers local teams, and we undertake hugely exciting work that is genuinely changing the world.

Our advanced portfolio of consulting, applications, business process, cloud, and infrastructure services will allow you to achieve great things by working with brilliant colleagues, and clients, on exciting projects.

Our inclusive work environment prioritises mutual respect, accountability, andcontinuous learning for all our people.

This approach fosters collaboration, well-being, growth, and agility, leading to a more diverse, innovative, and competitiveorganisation.

We are also proud to share that we have a range of Inclusion Networks such as: the Women’s Business Network, Cultural and Ethnicity Network, LGBTQ+ ll offer you: We offer a range of tailored benefits that support your physical, emotional, and financial wellbeing.

Our Learning and Development team ensure that there are continuous growth and development opportunities for our people.

We also offer the opportunity to have flexible work options.

You can find more information about NTT DATA UK

Site Reliability Engineering (SRE) / Observability Technical Lead employer: NTT DATA

At NTT DATA, we pride ourselves on being an excellent employer that fosters a culture of innovation and collaboration. Our commitment to employee growth is evident through continuous learning opportunities and a supportive environment that encourages creativity in the rapidly evolving field of AI/ML security. Located in a vibrant tech hub, we offer competitive benefits and a dynamic workplace where your contributions directly impact the future of secure AI systems.

NTT DATA

Contact Details:

NTT DATA Recruitment Team

We think you need these skills to ace Site Reliability Engineering (SRE) / Observability Technical Lead

Application Performance Monitoring (APM)
Infrastructure as Code (IaC)
Automation
Distributed Tracing
OpenTelemetry
Terraform
Cloud Platforms (AWS, GCP, Azure)