Mode of working
Hybrid/office based
Hybrid
If Hybrid, how many days are required in office?
3 days
Number of positions
1
The Role
We are seeking an experienced Senior HPC Engineer to manage, secure and maintain the organisation's scientific computing environment. The role will focus on RHEL-based HPC infrastructure, Slurm workload management and scientific application support, while working closely with research scientists and technical teams to deliver reliable, secure and high-performing computing services.
Your responsibilities
- Administer, patch, secure and maintain RHEL 7, 8 and 9 environments across HPC clusters and high-end workstations.
- Deploy, configure and manage Slurm, including scheduling, queues, partitions and fair-share policies.
- Monitor and optimise cluster health, resource utilisation, storage, networking and job throughput.
- Install and support scientific software, compilers, libraries and MPI environments.
- Work directly with scientists to understand computational requirements, optimise workloads and resolve application issues.
- Troubleshoot complex hardware, operating system, scheduler and application problems, including root-cause analysis.
- Manage incidents, problems and service requests through ServiceNow.
- Collaborate with networking, storage, security, DevOps teams and external vendors to deliver reliable HPC services.
Your Profile
Essential skills/knowledge/experience
- Minimum of 10 years' enterprise IT experience, including at least 3 to 5 years in an HPC or research-computing role.
- Extensive hands‑on administration and troubleshooting experience with RHEL 7, 8 and 9.
- Proven experience managing HPC clusters and deploying, configuring and operating Slurm.
- Experience supporting scientific or research applications in a Linux HPC environment.
- Strong troubleshooting skills across hardware, operating systems, schedulers and applications.
- Working knowledge of ServiceNow or an equivalent ITSM platform.
- Ability to work collaboratively with research scientists and translate computational requirements into practical technical solutions.
- Excellent communication and stakeholder‑management skills.
- Ability to work onsite for a minimum of three days per week and attend at short notice when physical‑system support is required.
Desirable skills/knowledge/experience
- Experience with Docker or other container technologies in an HPC environment.
- Knowledge of Ansible or similar configuration‑management and automation tools.
- Experience supporting GPU computing, CUDA and GPU‑accelerated workloads on RHEL.
- Understanding of MPI libraries, particularly OpenMPI or MPICH, and their integration with Slurm.
- Familiarity with InfiniBand, high‑speed Ethernet and HPC networking concepts.
- Experience with web server configuration and SSL certificate management.
- Red Hat certifications such as RHCSA or RHCE, or equivalent qualifications.
#J-18808-Ljbffr
Senior HPC/Linux Engineer - in Stevenage employer: Infoplus Technologies UK Ltd
As a Sybase DBA at our company, you will thrive in a dynamic work environment that prioritises innovation and collaboration. We offer competitive benefits, a strong focus on employee development, and opportunities for growth within the organisation, all while being located in a vibrant area that fosters both professional and personal enrichment. Join us to be part of a team that values your contributions and supports your career aspirations.
Contact Details:
Infoplus Technologies UK Ltd Recruitment Team