Senior ML Engineer: Inference & Latency Optimization in London

Senior ML Engineer: Inference & Latency Optimization in London

London Full-Time On-site
N

Nebius is seeking a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts to production deployment on our Applied AI team. You will improve latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality.

This hands-on role involves diagnosing complex serving problems, evaluating configurations, and delivering measurable production improvements in collaboration with kernel and platform engineers.

#J-18808-Ljbffr

Senior ML Engineer: Inference & Latency Optimization in London employer: Nebius Group

Nebius is an exceptional employer, offering a dynamic and innovative work environment in the heart of London. With a strong focus on employee wellbeing and professional growth, we provide competitive compensation, flexible working arrangements, and the opportunity to contribute to groundbreaking AI projects alongside a talented team. Our collaborative culture fosters creativity and ownership, making it a truly rewarding place to build your career.

N

Contact Details:

Nebius Group Recruitment Team