Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure

Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure

Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
Sentinel

At a Glance

  • Tasks: Design and operate cutting-edge GPU compute environments for AI research.
  • Company: Join a leading tech firm at the forefront of AI/ML innovation.
  • Benefits: Competitive salary, hybrid work model, and opportunities for professional growth.
  • Other info: Dynamic team environment with a focus on collaboration and innovation.
  • Why this job: Make a real impact in AI research while working with advanced technologies.
  • Qualifications: Experience with HPC/GPU clusters and cloud platforms is essential.

The predicted salary is between 63000 - 77000 Β£ per year.

Sentinel is recruiting for several senior/staff-level engineers to design, build and operate a hybrid GPU compute environment, combining on-prem HPC clusters with public cloud infrastructure for large-scale AI research workloads.

Responsibilities:

  • Build and operate high-performance GPU training/inference clusters, including scheduling, isolation and automated life cycle management
  • Design high-throughput data paths across compute and storage, including parallel filesystems (eg Lustre)
  • Benchmark and resolve performance bottlenecks across compute, network and orchestration layers
  • Implement observability, resilience and security controls for a compliance-conscious research environment
  • Work with research and applied ML teams to forecast GPU/storage capacity and streamline experimentation pipelines

Requirements:

  • Experience with HPC/GPU clusters, including a strong understanding of GPU architecture, high-speed networking and distributed training performance
  • Cloud platform experience (Azure, GCP, AWS or other)
  • Experience with containerisation/Kubernetes
  • Working knowledge of IaC and CI/CD (eg Terraform, Argo CD)

Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure employer: Sentinel

Sentinel is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the heart of Central London. With a strong commitment to employee growth, we provide ample opportunities for professional development and impactful contributions to health transformation initiatives. Join us to be part of a rapidly growing team where your expertise will shape the future of healthcare consultancy.

Sentinel

Contact Details:

Sentinel Recruitment Team

We think you need these skills to ace Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure

HPC/GPU Clusters
GPU Architecture
High-Speed Networking
Distributed Training Performance
Cloud Platform Experience (Azure, GCP, AWS)
Containerisation
Kubernetes