At a Glance
- Tasks: Design and operate cutting-edge GPU compute environments for AI research.
- Company: Join a leading tech firm at the forefront of AI/ML innovation.
- Benefits: Competitive salary, hybrid work model, and opportunities for professional growth.
- Other info: Dynamic team environment with a focus on collaboration and innovation.
- Why this job: Make a real impact in AI research while working with advanced technologies.
- Qualifications: Experience with HPC/GPU clusters and cloud platforms is essential.
The predicted salary is between 63000 - 77000 Β£ per year.
Sentinel is recruiting for several senior/staff-level engineers to design, build and operate a hybrid GPU compute environment, combining on-prem HPC clusters with public cloud infrastructure for large-scale AI research workloads.
Responsibilities:
- Build and operate high-performance GPU training/inference clusters, including scheduling, isolation and automated life cycle management
- Design high-throughput data paths across compute and storage, including parallel filesystems (eg Lustre)
- Benchmark and resolve performance bottlenecks across compute, network and orchestration layers
- Implement observability, resilience and security controls for a compliance-conscious research environment
- Work with research and applied ML teams to forecast GPU/storage capacity and streamline experimentation pipelines
Requirements:
- Experience with HPC/GPU clusters, including a strong understanding of GPU architecture, high-speed networking and distributed training performance
- Cloud platform experience (Azure, GCP, AWS or other)
- Experience with containerisation/Kubernetes
- Working knowledge of IaC and CI/CD (eg Terraform, Argo CD)
Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure employer: Sentinel
Sentinel is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the heart of Central London. With a strong commitment to employee growth, we provide ample opportunities for professional development and impactful contributions to health transformation initiatives. Join us to be part of a rapidly growing team where your expertise will shape the future of healthcare consultancy.
We think you need these skills to ace Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure
HPC/GPU Clusters
GPU Architecture
High-Speed Networking
Distributed Training Performance
Cloud Platform Experience (Azure, GCP, AWS)
Containerisation
Kubernetes