Senior AI Infrastructure Engineer - Scale Multi-GPU Training

Senior AI Infrastructure Engineer - Scale Multi-GPU Training

Full-Time 60000 - 84000 Β£ / year (est.) No working from home possible
L

At a Glance

  • Tasks: Architect and optimise multi-GPU training in AWS, managing cluster orchestration with Slurm and Kubernetes.
  • Company: Join a leading tech firm at the forefront of AI innovation in London.
  • Benefits: Competitive salary, flexible working hours, and opportunities for professional growth.
  • Other info: Dynamic team environment with a focus on collaboration and innovation.
  • Why this job: Be part of groundbreaking AI projects and shape the future of technology.
  • Qualifications: Deep expertise in PyTorch, transformer models, and production AI systems.

The predicted salary is between 60000 - 84000 Β£ per year.

Linux Recruit is seeking an experienced AI Infrastructure or MLOps Engineer to join our London lab.

You will architect and optimize distributed training across multiple GPUs and machines in AWS, eliminate bottlenecks in the data path, and manage cluster orchestration with Slurm and Kubernetes.

The role requires deep Py Torch expertise, familiarity with transformer models, and experience deploying production AI systems.

#J-18808-Ljbffr

Senior AI Infrastructure Engineer - Scale Multi-GPU Training employer: LinuxRecruit

At LinuxRecruit, we pride ourselves on being an excellent employer by fostering a collaborative and innovative work culture in the heart of London. Our team-oriented environment encourages personal growth and development, allowing you to tackle real-world challenges while directly engaging with clients. With attractive compensation packages and a commitment to high trust and low ego, we offer a unique opportunity for meaningful and rewarding employment in the rapidly evolving field of AI systems.

L

Contact Details:

LinuxRecruit Recruitment Team

We think you need these skills to ace Senior AI Infrastructure Engineer - Scale Multi-GPU Training

AI Infrastructure Engineering
MLOps
Distributed Training
Multi-GPU Systems
AWS
Data Path Optimization
Cluster Orchestration