HPC Engineer - Slurm Expertise

HPC Engineer - Slurm Expertise

Full-Time On-site
T

Job Title: HPC Engineer


Job Description


We are seeking an experienced Senior HPC Engineer with a core expertise in the Slurm workload manager. This role is pivotal in owning scheduling, job and queue design, and cluster policy management across our HPC estate. The primary mandate of this position is Slurm ownership.


Responsibilities



  • Own and manage the Slurm workload manager.

  • Design and implement job and queue scheduling strategies.

  • Develop and enforce cluster policies across the HPC estate.

  • Manage, secure, and maintain Red Hat Enterprise Linux (RHEL) infrastructure across versions 7, 8, and 9.

  • Provide support for scientific applications and facilitate collaboration with research scientists.


Essential Skills



  • Minimum of 10 years in enterprise IT with significant hands‑on experience in High-Performance Computing (HPC) environments.

  • Deep expertise in Slurm workload management.

  • Proficiency in managing RHEL infrastructure.

  • Strong knowledge of GPU/CUDA technologies.

  • Proven ability in root cause analysis.


Location


Stevenage, UK


Rate/Salary


600.00 GBP Daily

#J-18808-Ljbffr

HPC Engineer - Slurm Expertise employer: Teksystems

As an Azure Solution Architect with us, you'll thrive in a dynamic work culture that prioritises innovation and collaboration. We offer competitive benefits, including professional development opportunities to enhance your skills in cloud technologies, all while working in a vibrant location that fosters creativity and teamwork. Join us to be part of a forward-thinking company that values your contributions and supports your career growth.

T

Contact Details:

Teksystems Recruitment Team