Job summaryAn SRE engineer will apply engineering principles to remediate infrastructure and operational problems. The primary focus will be on automation and CI/CD; ensuring our services run reliably, are scalable, and perform optimally in production environments. The role will monitor and manage these aspects while taking responsibility for multiple cloud infrastructure services. Observability of systems will be key to prioritising the operational service improvements and performance improvements to meet and exceed SLOs (Service Level Objectives).Main duties of the jobWorking with the HPC & SRE Team to:Ensure services are stable, scalable, performant and automatedRespond to incidents, troubleshooting issues, and restoring services as quickly as possiblePrioritise operational service improvements to meet or increase SLO, minimising downtimeEnsure that effective monitoring/alerting is in place to proactively identify issues using tools and dashboards. Reducing times to respond to issues.Leverage automation to streamline tasks, reduce overhead on repeatable operations, reduce manual intervention and improve efficiency. Write code that is maintainable, clear, and conciseOptimise system performance using strong problem-solving skills to identify bottlenecks with an engineering mindsetEnsure systems can handle current and future workloads through automation and capacity planningContinuously improve services through observability, and identify ways to improve observability practicesFollow SRE principles. Guide and educate stakeholders to adopt implemented principlesProvide technical documentation for engineers. Providing training, where appropriateWorking closely with engineering and technology teams to improve operational processes, reduce manual tasks, ensure seamless collaboration/knowledge sharing, reduce risks and adapt to new ways of working.About usWe pride ourselves as being an employer of choice, where Everyone Matters promoting equality of opportunity to actively encourage applications from everyone, including groups currently underrepresented in our workforce.UKHSA ethos is to be an inclusive organisation for all our staff and stakeholders. To create, nurture and sustain an inclusive culture, where differences drive innovative solutions to meet the needs of our workforce and wider communities. We do this through celebrating and protecting differences by removing barriers and promoting equity and equality of opportunity for all.Please visit our careers site for more information posted23 September 2026Pay schemeOtherSalary
41,983 to 52,113 a year
per annum, pro rata (+MPS of up to 5,000 reviewed 31/3/27)
ContractPermanentWorking pattern
Full-time,
Part-time,
Job share,
Flexible working
Reference number919-NP- -EXTJob locationsCore HQ or Scientific CampusBirmingham, Chilton, Leeds, Liverpool, London, PortonE14 4PUUnited KingdomJob descriptionJob responsibilitiesWe are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to join our HPC & SRE engineering team. As an SRE, you will play a critical role in ensuring the stability, scalability, and performance of our services. You will combine software engineering and systems engineering to build, improve and run reliable, scalable production systems.Key Responsibilities:Service Reliability & PerformanceEnsure services are stable, scalable, and performant through engineering best practices and system design.Proactively identify and address system bottlenecks using advanced problem-solving and performance tuning techniques.Conduct capacity planning and implement solutions to ensure systems can support current and future workloads.Incident Response & TroubleshootingRespond swiftly to production incidents, ensuring minimal downtime and quick restoration of services.Perform root cause analysis and postmortems, implementing lessons learned to prevent recurrence.Monitoring, Alerting & ObservabilityContribute to the design and implementation of effective monitoring and alerting systems using tools and dashboards.Improve observability of services, ensuring issues are identified and addressed before impacting users.Continuously refine monitoring practices to reduce alert fatigue and improve response times.Automation & ToolingDevelop automation to eliminate manual, repetitive tasks and improve operational efficiency.Write clear, maintainable, and well-tested code to support automation efforts and system tooling.Drive initiatives to reduce operational toil and improve reliability through Infrastructure as Code (IaC).Service Level Objectives (SLOs) & Operational ImprovementsContribute to the definition, tracking, and continuous improvement of SLOs, SLIs, and error budgets.Identify and prioritize operational improvements that align with business goals and user experience.SRE Best Practices & AdvocacyHelping to evangelize SRE principles across the organization.Collaborate with stakeholders to integrate reliability practices into the development lifecycle.Collaboration & Knowledge SharingWork closely with software engineering, DevOps, and infrastructure teams to streamline deployment and operational workflows.Improve cross-functional collaboration and promote a culture of shared responsibility for service reliability.Documentation & TrainingMaintain accurate technical documentation, runbooks, and post-incident reports.Provide training and mentorship to engineering teams on best practices and tools.This list is not exhaustive .Essential Criteria:Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar roleCoding skills in programming/scripting languages such as Python, PowerShell or BashUnderstanding of Linux/Unix & Windows systems, networking, and distributed systemsExperience with observability tools (e.g., Prometheus, Grafana, Datadog) and alerting systemsUnderstanding of infrastructure automation (e.g., Terraform, Ansible, PowerShell, Helm)Excellent communication and collaboration skillsPossesses problem solving skills and the ability to respond to sudden unexpected demandsDesirable criteria:Experience with CI/CD pipelines, cloud platforms (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes)Experience with post-incident reviewsPrevious involvement in driving adoption of SRE practices across an organizationExperience delivering training or mentoring junior engineersSelection ProcessDetailsThis vacancy is using SuccessProfiles andwill assess your Behaviours, Experience and Technical skills.Stage 1: Application &SiftSuccess Profiles - GOV.UKYou willbe requiredto complete an application form. You will be assessed onthe listed 7 essentialcriteria,and this will be in the form ofa:Application form(Employer/ Activity history section on the application)1000 wordsupporting statement.This should outline how you consider your skills,experienceandknowledgeprovide evidence of your suitability for the role, with reference to the essential criteria.You will receive a joint score for your application form and statement. (The application form is the kind of information you would put into your C.V please beadvisedyou will not be able to upload your CV. Please complete the application form in as much detail as possible). Please do not email us your CV.Longlisting:In the event ofa large number ofapplicationswewilllonglistinto 3 piles of:Meets all essential criteriaMeets some essential criteriaMeets no essential criteriaIf used, the pile(s) Meets all essential criteria willproceedto shortlisting.Shortlisting:In the event ofa large number ofapplications wemay conductan initialsift, on the lead criteria of:Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar roleDesirable criteria maybeusedin the event ofa large number ofapplications/largeamountof successful candidates.If you are successful at this stage, you will progress to interview & assessment.Feedback will not be provided at this stage.Stage 2:InterviewSuccess Profiles - GOV.UKYou will be invited to a face to faceinterview at Canary Wharf. London.If face to face interviewsareplanned, in exceptional circumstances, we may be able to offer a remote interview.Behaviours and technical skills will be tested at interview through questioning and a presentation.There will be a Presentation required, the subject being:Automating a complex operational process.The Behaviours tested during the interview stage will be:Changing and improving. (Lead behaviour)Working together.Managing a quality service.Delivering at pace.Interviewdatesareto be confirmed.Once this job has closed, the job advert will no longer be available. You may want to save a copy for your records.LocationThis role is being offered as hybrid working based at one of our Core HQ's or Scientific Campuses.We offer great flexible working opportunities at UKHSA and operate using a hybrid working model where business needs allow. This provides us with greater flexibility about how and where we work, to get the best from our workforce. As a hybrid worker, you will be expected to spend a minimum of 60% of your contractual working hours (approximately 3 days a week pro rata, (averaged over a month) working at one of UKHSA's core HQs (Birmingham, Leeds, Liverpool, and London), or scientific campus sites (Colindale, Porton or Chilton).Our core HQ offices are modern and newly refurbished with excellent city centre transport links and benefit from co-location with other government departments such as the Department for Health and Social Care (DHSC).Security Clearance Level RequirementSuccessful candidates must pass a basic disclosure and barring security check before they can be appointed.Successful candidates must meet the security requirements before they can be appointed. The level of security needed is Security Check.For meaningful National Security Vetting checks to be carried out individuals need to have lived in the UK for a sufficient period of time. You should normally have been resident in the United Kingdom for the last 5 years as the role requires Security Check (SC) clearance. UKresidency less than the outlined periods may not necessarily bar you from gaining national security vetting and applicants should contact the Vacancy Holder/Recruiting Manager listed in the advert for further advice.Successful candidates must meet the security requirements before they can be appointed.Eligibility CriteriaExternal:This vacancy is open to all external applicants (anyone) from outside the CivilService as well as internal applicants.Salary InformationIf you are successful at interview, and are moving from another government department, NHS, or Local Authority, the relevant starting salary principles for level transfers or promotions will apply. Otherwise, roles are offered at the pay scale minimum for the grade, but in exceptional circumstances there may be flexibility if you are able to demonstrate you are already in receipt of an existing, higher salary. Pay increases are through the relevant annual pay award for the role and terms.Please be aware that the salary is based on the office location.Senior Executive Officer (SEO)41,983- 48,128 (National) (+MPS of up to 5,000 per annum, pro rata)44,148- 50,121 (Outer London) (+MPS of up to 5,000 per annum, pro rata)46,310- 52,113 (Inner London) (+MPS of up to 5,000 per annum, pro rata)Market Pay Supplement (MPS) is offered for this role of up to 5,000 per annum, pro rata. This is due to be reviewed 31/3/27. The amount offered in MPS will be dependant on the postholder's capability level. Each post holder would be assessed on entry to the civil service with the majority being placed on the lowest award unless clear evidence can be provided of higher capability level. The MPS bands are:Developing - up to 2,500 per annum, pro rataProficient - up to 3,500 per annum, pro rataAccomplished - up to 5,000 per annum, pro rata
Job description
Job responsibilitiesWe are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to join our HPC & SRE engineering team. As an SRE, you will play a critical role in ensuring the stability, scalability, and performance of our services. You will combine software engineering and systems engineering to build, improve and run reliable, scalable production systems.Key Responsibilities:Service Reliability & PerformanceEnsure services are stable, scalable, and performant through engineering best practices and system design.Proactively identify and address system bottlenecks using advanced problem-solving and performance tuning techniques.Conduct capacity planning and implement solutions to ensure systems can support current and future workloads.Incident Response & TroubleshootingRespond swiftly to production incidents, ensuring minimal downtime and quick restoration of services.Perform root cause analysis and postmortems, implementing lessons learned to prevent recurrence.Monitoring, Alerting & ObservabilityContribute to the design and implementation of effective monitoring and alerting systems using tools and dashboards.Improve observability of services, ensuring issues are identified and addressed before impacting users.Continuously refine monitoring practices to reduce alert fatigue and improve response times.Automation & ToolingDevelop automation to eliminate manual, repetitive tasks and improve operational efficiency.Write clear, maintainable, and well-tested code to support automation efforts and system tooling.Drive initiatives to reduce operational toil and improve reliability through Infrastructure as Code (IaC).Service Level Objectives (SLOs) & Operational ImprovementsContribute to the definition, tracking, and continuous improvement of SLOs, SLIs, and error budgets.Identify and prioritize operational improvements that align with business goals and user experience.SRE Best Practices & AdvocacyHelping to evangelize SRE principles across the organization.Collaborate with stakeholders to integrate reliability practices into the development lifecycle.Collaboration & Knowledge SharingWork closely with software engineering, DevOps, and infrastructure teams to streamline deployment and operational workflows.Improve cross-functional collaboration and promote a culture of shared responsibility for service reliability.Documentation & TrainingMaintain accurate technical documentation, runbooks, and post-incident reports.Provide training and mentorship to engineering teams on best practices and tools.This list is not exhaustive .Essential Criteria:Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar roleCoding skills in programming/scripting languages such as Python, PowerShell or BashUnderstanding of Linux/Unix & Windows systems, networking, and distributed systemsExperience with observability tools (e.g., Prometheus, Grafana, Datadog) and alerting systemsUnderstanding of infrastructure automation (e.g., Terraform, Ansible, PowerShell, Helm)Excellent communication and collaboration skillsPossesses problem solving skills and the ability to respond to sudden unexpected demandsDesirable criteria:Experience with CI/CD pipelines, cloud platforms (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes)Experience with post-incident reviewsPrevious involvement in driving adoption of SRE practices across an organizationExperience delivering training or mentoring junior engineersSelection ProcessDetailsThis vacancy is using SuccessProfiles andwill assess your Behaviours, Experience and Technical skills.Stage 1: Application &SiftSuccess Profiles - GOV.UKYou willbe requiredto complete an application form. You will be assessed onthe listed 7 essentialcriteria,and this will be in the form ofa:Application form(Employer/ Activity history section on the application)1000 wordsupporting statement.This should outline how you consider your skills,experienceandknowledgeprovide evidence of your suitability for the role, with reference to the essential criteria.You will receive a joint score for your application form and statement. (The application form is the kind of information you would put into your C.V please beadvisedyou will not be able to upload your CV. Please complete the application form in as much detail as possible). Please do not email us your CV.Longlisting:In the event ofa large number ofapplicationswewilllonglistinto 3 piles of:Meets all essential criteriaMeets some essential criteriaMeets no essential criteriaIf used, the pile(s) Meets all essential criteria willproceedto shortlisting.Shortlisting:In the event ofa large number ofapplications wemay conductan initialsift, on the lead criteria of:Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar roleDesirable criteria maybeusedin the event ofa large number ofapplications/largeamountof successful candidates.If you are successful at this stage, you will progress to interview & assessment.Feedback will not be provided at this stage.Stage 2:InterviewSuccess Profiles - GOV.UKYou will be invited to a face to faceinterview at Canary Wharf. London.If face to face interviewsareplanned, in exceptional circumstances, we may be able to offer a remote interview.Behaviours and technical skills will be tested at interview through questioning and a presentation.There will be a Presentation required, the subject being:Automating a complex operational process.The Behaviours tested during the interview stage will be:Changing and improving. (Lead behaviour)Working together.Managing a quality service.Delivering at pace.Interviewdatesareto be confirmed.Once this job has closed, the job advert will no longer be available. You may want to save a copy for your records.LocationThis role is being offered as hybrid working based at one of our Core HQ's or Scientific Campuses.We offer great flexible working opportunities at UKHSA and operate using a hybrid working model where business needs allow. This provides us with greater flexibility about how and where we work, to get the best from our workforce. As a hybrid worker, you will be expected to spend a minimum of 60% of your contractual working hours (approximately 3 days a week pro rata, (averaged over a month) working at one of UKHSA's core HQs (Birmingham, Leeds, Liverpool, and London), or scientific campus sites (Colindale, Porton or Chilton).Our core HQ offices are modern and newly refurbished with excellent city centre transport links and benefit from co-location with other government departments such as the Department for Health and Social Care (DHSC).Security Clearance Level RequirementSuccessful candidates must pass a basic disclosure and barring security check before they can be appointed.Successful candidates must meet the security requirements before they can be appointed. The level of security needed is Security Check.For meaningful National Security Vetting checks to be carried out individuals need to have lived in the UK for a sufficient period of time. You should normally have been resident in the United Kingdom for the last 5 years as the role requires Security Check (SC) clearance. UKresidency less than the outlined periods may not necessarily bar you from gaining national security vetting and applicants should contact the Vacancy Holder/Recruiting Manager listed in the advert for further advice.Successful candidates must meet the security requirements before they can be appointed.Eligibility CriteriaExternal:This vacancy is open to all external applicants (anyone) from outside the CivilService as well as internal applicants.Salary InformationIf you are successful at interview, and are moving from another government department, NHS, or Local Authority, the relevant starting salary principles for level transfers or promotions will apply. Otherwise, roles are offered at the pay scale minimum for the grade, but in exceptional circumstances there may be flexibility if you are able to demonstrate you are already in receipt of an existing, higher salary. Pay increases are through the relevant annual pay award for the role and terms.Please be aware that the salary is based on the office location.Senior Executive Officer (SEO)41,983- 48,128 (National) (+MPS of up to 5,000 per annum, pro rata)44,148- 50,121 (Outer London) (+MPS of up to 5,000 per annum, pro rata)46,310- 52,113 (Inner London) (+MPS of up to 5,000 per annum, pro rata)Market Pay Supplement (MPS) is offered for this role of up to 5,000 per annum, pro rata. This is due to be reviewed 31/3/27. The amount offered in MPS will be dependant on the postholder's capability level. Each post holder would be assessed on entry to the civil service with the majority being placed on the lowest award unless clear evidence can be provided of higher capability level. The MPS bands are:Developing - up to 2,500 per annum, pro rataProficient - up to 3,500 per annum, pro rataAccomplished - up to 5,000 per annum, pro rataPerson Specification
Application form and supporting statement
EssentialApplication form and supporting statement
Behaviours
EssentialChanging and improving. (Lead behaviour)Working together.Managing a quality service.Delivering at pace.
Technical questions
EssentialTechnical questions
Presentation
EssentialPresentation
Person Specification
Application form and supporting statement
EssentialApplication form and supporting statement
Behaviours
EssentialChanging and improving. (Lead behaviour)Working together.Managing a quality service.Delivering at pace.
Technical questions
EssentialTechnical questions
Presentation
EssentialPresentationDisclosure and Barring Service CheckThis post is subject to the Rehabilitation of Offenders Act (Exceptions Order) 1975 and as such it will be necessary for a submission for Disclosure to be made to the Disclosure and Barring Service (formerly known as CRB) to check for any previous criminal convictions.Certificate of SponsorshipApplications from job seekers who require current Skilled worker sponsorship to work in the UK are welcome and will be considered alongside all other applications. For further information visit the UK Visas and Immigration website (Opens in a new tab).From 6 April 2017, skilled worker applicants, applying for entry clearance into the UK, have had to present a criminal record certificate from each country they have resided continuously or cumulatively for 12 months or more in the past 10 years. Adult dependants (over 18 years old) are also subject to this requirement. Guidance can be found here Criminal records checks for overseas applicants (Opens in a new tab).
Additional information
Disclosure and Barring Service CheckThis post is subject to the Rehabilitation of Offenders Act (Exceptions Order) 1975 and as such it will be necessary for a submission for Disclosure to be made to the Disclosure and Barring Service (formerly known as CRB) to check for any previous criminal convictions.Certificate of SponsorshipApplications from job seekers who require current Skilled worker sponsorship to work in the UK are welcome and will be considered alongside all other applications. For further information visit the UK Visas and Immigration website (Opens in a new tab).From 6 April 2017, skilled worker applicants, applying for entry clearance into the UK, have had to present a criminal record certificate from each country they have resided continuously or cumulatively for 12 months or more in the past 10 years. Adult dependants (over 18 years old) are also subject to this requirement. Guidance can be found here Criminal records checks for overseas applicants (Opens in a new tab).Employer detailsEmployer nameUK Health Security AgencyAddressCore HQ or Scientific CampusBirmingham, Chilton, Leeds, Liverpool, London, PortonE14 4PUUnited KingdomEmployer's websiteEmployer detailsEmployer nameUK Health Security AgencyAddressCore HQ or Scientific CampusBirmingham, Chilton, Leeds, Liverpool, London, PortonE14 4PUUnited KingdomEmployer's website
Senior Specialist Engineer (Specialist Site Reliability Engineer SRE) in London employer: National Health Service
Betsi Cadwaladr University Health Board is an exceptional employer, offering a supportive and collaborative work environment for healthcare professionals in North Wales. With a commitment to employee development and a focus on compassionate care, staff have access to continuous professional development opportunities and the chance to make a meaningful impact in both acute and community paediatrics. The Health Board's integrated approach ensures that employees are part of a dynamic team dedicated to improving health outcomes for the local population.