Senior Site Reliability Engineer – Python in Cardiff

Senior Site Reliability Engineer – Python in Cardiff

Cardiff Full-Time No working from home possible
A

Senior Site Reliability Engineer – Python

  • Company: Government Recruitment Service
  • Salary: £63,824 - £83,778
  • Hours: Full-time
  • Location: Cardiff
  • Job type: Permanent
  • Posting date: 25 Aug 2026
  • Closing date: 14 Sept 2026

If you would like to find out more about the role, the SRE team and what it’s like to work at BIST, we are holding a Hiring Manager Q&A session for this role where you can virtually 'meet the team' on Tuesday 8th Septemberat 12:30pm.

About us

The Department for Business, Innovation, Science and Trade (BIST) helps businesses invest, innovate, export and grow across the UK. Digital, Data and Technology (DDaT) helps deliver this by designing, building and running the services, platforms and technology that businesses and colleagues rely on every day.

Our teams work on a wide range of services and technologies, including:

  • digital services that help businesses access government support, guidance and opportunities
  • data and AI products that support decision making and enable innovation
  • internal tools and services that help colleagues work more effectively
  • secure technology and platforms that support critical government services
  • cyber security services that protect systems, data and users

By joining DDaT, you'll work on services used by businesses, colleagues and citizens across the UK. You'll be part of multidisciplinary teams focused on improving services, solving complex problems and delivering better outcomes for users.

We are committed to creating an inclusive and supportive workplace. In recognition of this, we were named Best Public Sector Employer at the Women in Tech Employer Awards 2025!

About the role

As a Senior Site Reliability Engineer (DevOps), you will play a key role in designing,buildingand running reliable,scalableand secure platform services that underpin criticalBISTdigital products. Working in multidisciplinary, agile teams,you’llhelp ensure development teams have the tools and support they need,from observability andmonitoringthrough to CI/CD pipelinesso services are resilient,performantand centred around user needs.You’llchampion good engineering practices, helping teams use data and insight to continuously improve how services are delivered andoperated.

This is a hands-on role whereyou’llspend much of your time building and improving platform capabilities directly.You’llwork closely with product managers,architectsand engineers to improve reliability and reduce operational burden, while also supporting and mentoring others across the engineering community.You’llcontribute to building and scaling our global platform, support live services through an on-call rota, and play an active role in initiatives such as enhancing observability and streamlining deployment processes to improve service quality and delivery outcomes.

You will:

  • Build andmaintainshared service products enabling developers to be more efficient.
  • Write clean,maintainableand well-tested code to support platform and tooling development.
  • Work closely with development teams to provide and improve platform tooling, including monitoring, logging, metrics,dashboardsand CI/CD pipelines.
  • Build,maintainand improve reliable,secureand scalable cloud-based infrastructure using infrastructure-as-code approaches.
  • Support teams to adopt SRE practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs)and error budgets.
  • Contribute to observability across services, helping teams better understand performance,reliabilityand user impact.
  • Develop and improve CI/CD pipelines to enable safe,frequentand low-risk delivery of changes.
  • Collaborate with product,deliveryand architecture colleagues to ensure platform services meet user and business needs.
  • Support live service operations, including incident response, troubleshooting and problem management, with a focus on learning and continuous improvement.
  • Share knowledge and provide coaching and mentoring to colleagues, contributing to a supportive and inclusive engineering community.
  • Contribute to improving security,resilienceand compliance practices within the platform.

What tech will you be using?

  • Python and Django framework
  • PostgreSQL as a service (Amazon RDS)
  • GitHub Actionsand AWSCodePipelines/CodeBuild
  • Terraform
  • Docker, Elastic Container Service (ECS)and Elastic Container Registry (ECR)

Proud member of the Disability Confident employer scheme

About Disability Confident

Disability Confident is a government scheme. It encourages employers to recruit and retain disabled people and those with long term health conditions.

Related jobs

SRE Squad Lead

£63,824 to £83,778 per year

Government Recruitment Service

Cardiff

Permanent Full time

Head of User Research

£74,639 to £96,980 per year

Government Recruitment Service

Cardiff

Permanent Full time

Data Engineer

United Welsh

Caerphilly

Permanent Full time Browse more jobs

#J-18808-Ljbffr

Senior Site Reliability Engineer – Python in Cardiff employer: Allscreens Nationwide Ltd

Allscreens Nationwide Ltd is an exceptional employer, offering a supportive and inclusive work culture that prioritises the wellbeing of both employees and the individuals they serve. With a strong focus on professional development, team leaders are encouraged to grow their skills in a dynamic environment while making a meaningful impact in the lives of adults with learning disabilities. Located in Leicester, the company provides a unique opportunity to be part of a dedicated team committed to delivering high-quality, person-centred care.

A

Contact Details:

Allscreens Nationwide Ltd Recruitment Team