AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops in Gloucester

AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops in Gloucester

Gloucester Full-Time No working from home possible
R

Radiant is seeking a senior Site Reliability Engineer in the UK to design, deploy, and operate scalable AI-native infrastructure. You will own Kubernetes clusters, tune Linux and I/O, and drive automation across the platform.

You will champion ITSM practices, maintain Prometheus/Grafana monitoring, and participate in 24x7 on-call support. Mentoring and cross-training with Platform SRE and HPC teams are key parts of the role.

#J-18808-Ljbffr

AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops in Gloucester employer: Radiant

Radiant is an exceptional employer for those looking to make a significant impact in the AI infrastructure space. With a startup culture that values adaptability and hands-on involvement, employees are encouraged to build and innovate rather than follow pre-existing protocols. The company offers unique opportunities for professional growth, a collaborative work environment, and the chance to set industry standards in safety and quality management while working on cutting-edge projects.

R

Contact Details:

Radiant Recruitment Team