Site Reliability Engineer in London

Site Reliability Engineer in London

London Full-Time No working from home possible
W

Role summary: Applies software engineering practice to infrastructure and operations, with

accountability for the availability, performance, and security of business-critical internal tools and

services.

Scope of the role

β€’ Owns availability, latency, and performance against defined SLOs and error budgets, prioritised

by business impact.

β€’ Implements infrastructure as code and automated deployment pipelines that make releases

repeatable and rollbacks rapid.

β€’ Establishes observability (telemetry, actionable alerting, and service health dashboards) with

appropriate handling of confidential data.

β€’ Maintains reliable integrations with upstream internal systems, including authentication and

access provisioning.

β€’ Leads incident response, communicates status to affected users, and runs post-incident

reviews with tracked remediation.

β€’ Eliminates operational toil through automation, treating recurring manual intervention as a

defect.

Requirements

β€’ 5+ years in SRE, DevOps, or infrastructure engineering supporting production services.

β€’ Strong operational programming and automation capability:5 Python, Go, and shell scripting

(Bash/zsh).

β€’ Deep knowledge of Linux and macOS internals, networking fundamentals (TCP/IP, DNS, TLS,

load balancing), and container orchestration at scale (Kubernetes, Docker).

β€’ Practical experience with infrastructure as code (Terraform, Ansible) and observability platforms

(Prometheus, Grafana, OpenTelemetry, or equivalent).

β€’ Working knowledge of CI/CD integration, artifact management, and staged release strategies.

β€’ Experience operating web applications and API services, including database administration,

backup and recovery, and performance tuning.

β€’ Familiarity with enterprise identity and access management (SSO, OAuth/SAML, RBAC) and

handling confidential data under internal security requirements.

β€’ Experience with AI/ML or GPU-backed workloads is advantageous.

β€’ Composure during incidents, a bias toward durable remediation, and strong written

communication for runbooks and reviews.

W

Contact Details:

Wipro Europe Recruitment Team