Staff AI Cloud SRE: GPU Compute & ML Platform

Staff AI Cloud SRE: GPU Compute & ML Platform

Full-Time No working from home possible
Talanto

At a Glance

  • Tasks: Build and scale the reliability of our AI cloud platform with a focus on GPU Compute.
  • Company: Wayve, a pioneering tech company in AI and cloud solutions.
  • Benefits: Hybrid work model, competitive salary, and opportunities for professional growth.
  • Other info: Dynamic team environment with a strong focus on innovation and collaboration.
  • Why this job: Join us to shape the future of AI infrastructure and make a real impact.
  • Qualifications: Experience in SRE practices and managing large GPU-backed clusters.

Wayve is seeking a Staff Cloud Site Reliability Engineer (AI) to build and scale the reliability foundations of our AI cloud platform, including the Model Development Platform and GPU Compute, with a focus on resilient, efficient, and scalable model development infrastructure.

This London-based role offers hybrid work (2 days in the office) and requires owning SRE practices, on-call rotations, monitoring, and automation across large GPU-backed clusters to accelerate training and deployment.

Staff AI Cloud SRE: GPU Compute & ML Platform employer: Talanto

Almedia is an exceptional employer that fosters a dynamic and innovative work culture, perfect for those passionate about machine learning and real-time personalization. With a strong emphasis on employee growth, you will have the opportunity to mentor fellow engineers while collaborating with cross-functional teams in the vibrant city of London. The hybrid working arrangements and commitment to impactful projects make Almedia a rewarding place to advance your career in AdTech.

Talanto

Contact Details:

Talanto Recruitment Team

We think you need these skills to ace Staff AI Cloud SRE: GPU Compute & ML Platform

Site Reliability Engineering (SRE)
Cloud Infrastructure Management
GPU Compute
Model Development Platform
Monitoring and Automation
Scalability
Resilience Engineering