Salary: Β£41,500 - 69,500 per year
Requirements:
- We require 7+ years in SRE, DevOps, or production engineering, with 3+ years in a senior or lead capacity.
- We expect proven experience improving availability, reducing MTTR, and implementing self-healing at scale.
- We require experience managing or automating database operations in enterprise environments.
- We prefer relevant certifications such as CKA, AWS DevOps Professional, Azure DevOps Expert, or SRE Foundation.
- We require a bachelors degree in Computer Science, Engineering, or a related field, or equivalent experience.
- We need expert-level observability experience with tools such as Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace.
- We need strong experience with AIOps and ML-driven monitoring, such as PagerDuty, Moogsoft, BigPanda, or custom ML pipelines.
- We require deep knowledge of FMEA, fault tree analysis, and chaos engineering tools such as Gremlin, LitmusChaos, or Chaos Monkey.
- We require strong SQL skills plus experience with automated database provisioning, migration tools such as Liquibase or Flyway, and DB release pipelines.
- We need proficiency in automation and scripting with Python, Go, or Bash, including self-healing runbooks.
- We require infrastructure knowledge across Kubernetes, cloud platforms such as AWS, Azure, or GCP, networking, and storage systems.
- We need CI/CD and release engineering experience with Jenkins, GitLab CI, Spinnaker, or ArgoCD for integrated DB and application releases.
- We expect experience with FinOps principles, resource optimisation, and cloud spend analysis.
- We require strong leadership, mentoring, stakeholder management, communication, analytical, and problem-solving skills.
- We need the ability to manage competing priorities across multiple workstreams simultaneously.
- Desirable: published work or conference talks on SRE, observability, or chaos engineering.
- Desirable: experience with service mesh such as Istio and distributed tracing at scale.
- Desirable: background in financial services or regulated industry SRE practices.
Responsibilities:
- We lead the Site Reliability Engineering practice and drive the transformation from reactive operations to proactive, engineering-led reliability.
- We define and enforce non-functional requirements for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis.
- We design and implement self-healing automation for known failure patterns to reduce human intervention and on-call burden by 50%+.
- We build comprehensive observability stacks with metrics, logs, and traces, including ML-driven anomaly detection and AIOps capabilities.
- We lead automated incident management, including detection, triage, escalation, remediation, and post-incident review automation.
- We drive database automation, including automated provisioning, release management, backup and restore, and operational request workflows.
- We define and track SLIs, SLOs, and error budgets across critical services to balance reliability with feature velocity.
- We conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes.
- We mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement.
- We collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines.
Technologies:
- AWS
- OpenSearch
- ArgoCD
- Azure
- Bash
- CI/CD
- Cloud
- Datadog
- DevOps
- Dynatrace
- ELK
- Flyway
- GCP
- GitLab
- Grafana
- Incident Management
- Istio
- Support
- Jenkins
- Kubernetes
- Liquibase
- PagerDuty
- Prometheus
- Python
- SQL
- Architect
More:
We are Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our worlds potential. We are a people-centric team focused on powering good and helping future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. We turn organizations into data-driven leaders through engineering excellence and innovation, and we value diverse perspectives, inclusion, mutual respect, and merit-based systems. We offer industry-leading benefits and support for holistic health and wellbeing, flexible arrangements where role and location allow, and a sense of belonging, autonomy, freedom, and ownership. This full-time role is based in London, United Kingdom.
last updated 37 week of 2026
#J-18808-Ljbffr
SRE Architect employer: Hitachi
At Hitachi Energy, we pride ourselves on being an exceptional employer, offering a dynamic work environment in Birmingham that fosters collaboration and innovation. Our commitment to employee growth is evident through continuous training opportunities and a culture that values diversity and inclusion, ensuring every team member can thrive while contributing to a sustainable energy future. Join us to be part of a global team dedicated to making a meaningful impact in the energy sector.