Site Reliability Engineer (SRE) – Cloud Platforms in London

Site Reliability Engineer (SRE) – Cloud Platforms in London

London Full-Time No working from home possible
T

# Site Reliability Engineer (SRE) – Cloud Platforms,March 11, 2026### Job Description**Location:** London, UK **Work Model:** On-site **Role Type:** Full-TimeWe are looking for a **Site Reliability Engineer (SRE)** with strong experience in cloud infrastructure and production reliability to join our client’s on-site team in London.This role focuses on ensuring the reliability, scalability, and performance of mission-critical systems. You will work closely with software engineering and platform teams to build resilient infrastructure, automate operational processes, and improve system observability across production environments.---### **What You’ll Do*** Design and implement reliability strategies for high-availability production systems* Monitor system health, performance, and uptime across cloud infrastructure* Build automation to reduce manual operations and improve system reliability* Develop and maintain observability systems including logging, metrics, and tracing* Manage incident response processes and perform root cause analysis for production issues* Improve system resilience through capacity planning, performance optimisation, and fault tolerance* Collaborate with engineering teams to integrate reliability practices into the software development lifecycle* Implement infrastructure automation using Infrastructure as Code---### **What We’re Looking For**#### **Required Skills & Experience*** Strong experience operating production systems in cloud environments such as **Amazon Web Services**, **Google Cloud**, or **Microsoft Azure*** Experience with container orchestration platforms such as **Kubernetes*** Strong experience with monitoring and observability tools such as **Prometheus** and **Grafana*** Proficiency in scripting or programming languages such as Python, Go, or Bash* Experience implementing Infrastructure as Code with tools such as Terraform* Strong understanding of Linux systems, networking, and distributed systems---#### **Nice to Have*** Experience with CI/CD pipelines using platforms such as **GitHub** Actions or **GitLab*** Familiarity with incident management frameworks and reliability engineering practices (SLIs, SLOs, error budgets)* Experience supporting microservices architectures and high-scale systems* Knowledge of distributed tracing and performance monitoring---**Location:** London, UK **Work Model:** On-site **Role Type:** Full-TimeLocation,Experience levelMid–Senior level## Work Location
#J-18808-Ljbffr

Site Reliability Engineer (SRE) – Cloud Platforms in London employer: Talenzon group

At Talenzon group, we pride ourselves on fostering a collaborative and innovative work culture that empowers our employees to excel in their roles. As a Senior MLOps Engineer in our London office, you will benefit from a dynamic environment that encourages professional growth through hands-on experience with cutting-edge technologies like AWS and SageMaker. Our commitment to employee development, coupled with the vibrant city of London as your workplace, makes Talenzon an exceptional employer for those seeking meaningful and rewarding careers in AI and machine learning.

T

Contact Details:

Talenzon group Recruitment Team