Principal Site Reliability Engineer in London

Principal Site Reliability Engineer in London

London Full-Time On-site
G

Genomics England is a global leader in enabling genomic medicine and research, focused on creating a world where everyone benefits from genomic healthcare. Building on the 100,000 Genomes Project, we support the NHS’s world-first national whole genome sequencing service and run the growing National Genomic Research Library, alongside delivering numerous major genomics initiatives. By connecting research and clinical care at national scale, we enable immediate healthcare benefits and advances for the future.
Our mission is to provide the evidence and digital systems so that by 2035 genomics could play a role in up to half of all healthcare interactions, whilst securing the UK’s position as the best place to discover, prove and benefit from genomic innovations.
We are accelerating our impact and working with patients, doctors, scientists, government and industry to improve genomic testing, and help researchers access the health data and technology they need to make new medical discoveries and create more effective, targeted medicines for everybody.
Behind the Healthcare and Research outcomes, Genomics England delivers through designing, developing and operating complex healthcare software systems.
We're on the cusp of big changes with the real prospect of genomics becoming the fabric of everyday healthcare through the lifetime - from birth to old age.
Job Description
Principal Site Reliability Engineer
As Principal Site Reliability Engineer, you will help establish and grow Genomics England's organisation-wide SRE capability.
You will be responsible for:
Building and leading a small team of Site Reliability Engineers (initially 3 people)
Identifying the highest-value opportunities to improve reliability across our platforms
Defining standards, patterns and tooling that help engineering teams to build and operate reliable services
Developing services and capabilities that reduce operational toil and improve engineering effectiveness
Partnering with product and engineering teams to embed SRE principles into their ways of working
Influencing engineering strategy, governance and technical direction across the organisation
Helping to build a culture of reliability, continuous improvement and operational resilience
The SRE team will work alongside product squads, helping them adopt reliability practices, improve operations and solve complex technical challenges, rather than operating services on their behalf.
In your first year you will establish the initial SRE team, identify the highest-priority reliability challenges across Genomics England, and define the standards, services and practices that will underpin our long-term approach to reliability.
The Principal SRE reports to the Director of Engineering within the Technology and Product Directorate.
About the Tech Stack
The new SRE team will support squads that run a variety of services: most of these are either user-facing web applications (React), backend APIs (Python), bioinformatics pipelines (NextFlow), or data ETL workflows (Prefect, Dremio). These services increasingly run in AWS, though there is still a significant on-premise presence, and they run in a mixture of compute environments, from ECS/Fargate to HPC clusters to (occasionally) Kubernetes.
Within the SDLC we have a standard toolchain which includes Terraform for infrastructure-as-code, GitLab for source code and CI/CD, Artifactory for software artefacts, and DataDog for observability.
We are working to become interoperable with the wider NHS via open standards like FHIR and GA4GH APIs and increasingly aiming to integrate with their own API Management platform.
About You
You are a hands-on engineering leader with deep experience of SRE and platform engineering. You combine strategic thinking with the ability to identify and prioritise the areas where reliability improvements will have the greatest impact.
You are a problem-solver who identifies risks, issues, gaps, and dependencies and brings people together to find solutions. You do this through your supportive, empathetic and collaborative behaviours, acting as a coach, mentor, guide or constructive questioner as the situation demands.
You are a great communicator, comfortable not just leading your own team, but also engaging across the engineering community and with non-technical stakeholders.
Essential Skills and Experience
While we recognise the value of relevant qualifications or certifications, we are primarily interested in your real-world experience:
Comprehensive knowledge of SRE principles and practices with significant experience of applying these to real-world situations
Excellent software engineering skills especially in the context of release automation and other toil‑eliminating activities (Python preferred, polyglot ideal)
Strong understanding of how architecture and other factors contribute to the overall resilience of systems
Extensive experience of platform engineering across CI/CD, Infrastructure as Code, operational monitoring and alerting, backup and recovery etc.
Experience in at least one major public cloud (AWS preferred but not essential)
Demonstrable ability to lead teams, develop people and coordinate work towards shared outcomes
Strong interpersonal skills with a temperament that builds trust and connection within and across squads through open, honest communication
Comfortable engaging responsively with teams both remotely and in person when required
Ability to navigate rapidly to effective solutions through engaged and inclusive listening, clarity of thought, clear documentation, and succinct presentation
These skills are not essential but if you have either of them, they may prove to be useful:
Background in healthcare or bioinformatics
Experience in regulated environments
If you’re an experienced Site Reliability Engineer leader, who thrives on working collaboratively to mature engineering practices, we’d love to hear from you. Join us at Genomics England and make a meaningful impact in the world of genomics.
Additional Information
Salary From: £107,000
Closing Date: Monday 12th October @ 23:00 (UK time)
Being an integral part of such a meaningful mission is extremely rewarding in itself, but in order to support our people, we’re continually improving our benefits package. We pride ourselves on investing in our people and supporting them to achieve their career goals, as well as offering a benefits package including:
Generous Leave: 30 days’ holiday, plus

Principal Site Reliability Engineer in London employer: Genomics England

Genomics England is an exceptional employer, offering a dynamic work environment that fosters collaboration and innovation in the field of healthcare technology. With generous leave policies, flexible working arrangements, and a strong commitment to employee development, we empower our team members to achieve their career aspirations while contributing to meaningful projects that impact lives. Our inclusive culture and focus on well-being ensure that every employee feels valued and supported, making Genomics England a truly rewarding place to work.

G

Contact Details:

Genomics England Recruitment Team