Radiant is seeking a senior Infrastructure Site Reliability Engineer to own and improve large‑scale GPU‑accelerated HPC infrastructure in a 24/7 production environment.
You will work across network, storage, virtualization and orchestration with hands‑on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance benchmarking.
This role champions observability, automation and on‑call reliability, shaping next‑gen HPC platforms within a globally distributed team.
#J-18808-Ljbffr
24/7 HPC Infra SRE for AI & GPU Compute in Gloucester employer: Radiant
Radiant is an exceptional employer for those looking to make a significant impact in the AI infrastructure space. With a startup culture that values adaptability and hands-on involvement, employees are encouraged to build and innovate rather than follow pre-existing protocols. The company offers unique opportunities for professional growth, a collaborative work environment, and the chance to set industry standards in safety and quality management while working on cutting-edge projects.