Overview
In this role you will apply software and systems engineering to build and operate reliable, scalable production systems. You will focus on automation and CI/CD to maintain service reliability and performance in production. You’ll monitor across cloud services, drive observability improvements, and align with SLOs to minimize downtime. You’ll collaborate with cross-functional teams to promote reliable, efficient operations and share guidance on SRE practices.
Pay / Benefits
- hybrid working model
- flexible working opportunities
- uk government salary framework with market pay supplement (MPS) up to 5000
- core HQ locations with modern facilities
- inclusive culture promoting equality
Responsibilities
- Ensure services are stable, scalable, and performant through engineering best practices and system design
- Identify and address bottlenecks with advanced problem-solving and performance tuning
- Plan capacity to support current and future workloads
- Respond to production incidents and restore services quickly
- Perform root cause analysis and postmortems to prevent recurrence
- Design and implement monitoring/alerting systems using dashboards and tools
- Improve observability and reduce alert fatigue
- Develop automation to eliminate manual tasks and improve efficiency
- Write clear, maintainable code and drive IaC initiatives
- Contribute to SLO/SLI definition and continuous improvement of operational practices
- Advocate SRE principles and integrate reliability into development lifecycle
- Create and maintain technical documentation and provide training when appropriate
- Collaborate with software engineering, DevOps, and infrastructure teams to streamline deployment and operations
- Promote a culture of shared responsibility for service reliability
Key requirements
- Experience as Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar
- Coding skills in Python, PowerShell or Bash
- Understanding of Linux/Unix & Windows systems, networking, and distributed systems
- Experience with observability tools (Prometheus, Grafana, Datadog) and alerting
- Understanding of infrastructure automation (Terraform, Ansible, PowerShell, Helm)
- Excellent communication and collaboration skills
- Problem-solving ability to respond to sudden demands
- communication
- collaboration
- problem solving
- Python
- PowerShell
- Bash
Senior Specialist Engineer (Specialist Site Reliability Engineer SRE) in Chilton employer: National Health Service
Betsi Cadwaladr University Health Board is an exceptional employer, offering a supportive and collaborative work environment for healthcare professionals in North Wales. With a commitment to employee development and a focus on compassionate care, staff have access to continuous professional development opportunities and the chance to make a meaningful impact in both acute and community paediatrics. The Health Board's integrated approach ensures that employees are part of a dynamic team dedicated to improving health outcomes for the local population.