Senior GPU HPC Engineer: InfiniBand & KVM Optimization

Senior GPU HPC Engineer: InfiniBand & KVM Optimization

Full-Time 60000 - 80000 Β£ / year (est.) No working from home possible
Nebius

At a Glance

  • Tasks: Optimise GPU clusters and InfiniBand networks for high-performance computing.
  • Company: Join Nebius, a leader in hyperscaler technology with an innovative culture.
  • Benefits: Enjoy competitive pay, flexible work options, and opportunities for professional growth.
  • Other info: Fast-paced environment with excellent career advancement potential.
  • Why this job: Make a real impact on cutting-edge AI workloads across Europe and the UK.
  • Qualifications: Experience in HPC, GPU optimisation, and strong troubleshooting skills.

The predicted salary is between 60000 - 80000 Β£ per year.

Nebius is seeking a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will optimize GPU clusters, InfiniBand networks and the KVM/QEMU stack, collaborating with hardware virtualization and device emulation teams to ensure top performance and security in multi-GPU HPC environments.

The role emphasizes troubleshooting, integration of new hardware, and automation to sustain high-throughput AI workloads across Europe and the UK, with a fast-moving, innovative culture.

Senior GPU HPC Engineer: InfiniBand & KVM Optimization employer: Nebius

Nebius is an exceptional employer that fosters a dynamic and innovative work culture, particularly for those in the Remote Field CTO role. With flexible working options in London or remotely, employees benefit from a supportive environment that encourages professional growth and collaboration, while also having the unique opportunity to shape cutting-edge cloud services in the retail sector.

Nebius

Contact Details:

Nebius Recruitment Team

We think you need these skills to ace Senior GPU HPC Engineer: InfiniBand & KVM Optimization

GPU Cluster Optimization
InfiniBand Networking
KVM/QEMU Stack Management
Troubleshooting Skills
Hardware Integration
Automation Techniques
High-Throughput AI Workload Management