At a Glance
- Tasks: Lead the SRE practice, driving proactive reliability and automation.
- Company: Dynamic tech company focused on innovation and collaboration.
- Benefits: Competitive salary, flexible work options, and professional growth opportunities.
- Other info: Join a diverse team committed to continuous improvement and cutting-edge technology.
- Why this job: Make a real impact by enhancing system reliability and mentoring future engineers.
- Qualifications: 7+ years in SRE or DevOps with strong leadership skills.
The predicted salary is between 60000 - 80000 £ per year.
Overview
Lead the Site Reliability Engineering (SRE) practice, driving the transformation from reactive operations to proactive, engineering‑led reliability.
Own the definition and enforcement of non‑functional requirements (NFRs) using FMEA‑based resiliency frameworks, and champion observability, self‑healing automation, automated incident management, and database operations automation.
Key Responsibilities
- Define and enforce NFRs for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA‑based failure analysis.
- Design and implement self‑healing automation for known failure patterns, reducing human intervention and on‑call burden by 50%+.
- Build comprehensive observability stacks (metrics, logs, traces) with ML‑driven anomaly detection and AIOps capabilities.
- Lead automated incident management: detection, triage, escalation, remediation, and post‑incident review automation.
- Drive database automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations.
- Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity.
- Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes.
- Mentor two SRE engineers, establish engineering standards, and build a culture of reliability and continuous improvement.
- Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines.
- Technical Skills & Expertise
- Expert‑level observability: Prometheus, Grafana, ELK/Open Search, Jaeger/Zipkin, Datadog, or Dynatrace.
- Strong experience with AIOps and ML‑driven monitoring: Pager Duty, Moogsoft, Big Panda, or custom ML pipelines.
- Deep knowledge of FMEA, fault‑tree analysis, and chaos engineering tools (Gremlin, Litmus Chaos, Chaos Monkey).
- Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines.
- Proficiency in automation and scripting: Python, Go, Bash, with experience building self‑healing runbooks.
- Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems.
- CI/CD and release engineering: Jenkins, Git Lab CI, Spinnaker, Argo CD for integrated DB and application releases.
- Cost management: experience with Fin Ops principles, resource optimisation, and cloud spend analysis.
- Soft Skills & Competencies
- Strong leadership and mentoring ability—coaches and develops junior engineers.
- Excellent stakeholder management and communication skills across all levels.
- Strategic thinker who balances technical depth with business outcomes.
- Proven ability to drive change, influence without authority, and build consensus.
- Strong analytical and problem‑solving mindset with attention to detail.
- Ability to manage competing priorities across multiple workstreams simultaneously.
Qualifications & Experience
- 7+ years in SRE, Dev Ops, or production engineering with 3+ years in a senior or lead capacity.
- Proven track record of improving availability, reducing MTTR, and implementing self‑healing at scale.
- Experience managing or automating database operations in enterprise environments.
- Relevant certifications preferred: CKA, AWS Dev Ops Professional, Azure Dev Ops Expert, SRE Foundation.
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience).
Desirable / Nice to Have
- Published work or conference talks on SRE, observability, or chaos engineering.
- Experience with service mesh (Istio) and distributed tracing at scale.
- Background in financial services or regulated industry.
EEO Statement
We’re an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic.
Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
#J-18808-Ljbffr
SRE Architect (68019) employer: Hitachi Automotive Systems Americas, Inc.
Hitachi Rail UK Limited is an exceptional employer, offering a competitive salary and a comprehensive benefits package that includes private medical insurance and a generous pension scheme. Located in Bristol, the company fosters a supportive work culture that values diversity and inclusion, providing employees with opportunities for professional growth and development while ensuring a safe and efficient working environment for all team members.
Contact Details:
Hitachi Automotive Systems Americas, Inc. Recruitment Team