Nebius is seeking a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts to production deployment on our Applied AI team. You will improve latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality.
This hands-on role involves diagnosing complex serving problems, evaluating configurations, and delivering measurable production improvements in collaboration with kernel and platform engineers.
#J-18808-Ljbffr
Senior ML Engineer: Inference & Latency Optimization employer: Nebius Group
Nebius is an exceptional employer, offering a dynamic and innovative work environment in the heart of London. With a strong focus on employee wellbeing and professional growth, we provide competitive compensation, flexible working arrangements, and the opportunity to contribute to groundbreaking AI projects alongside a talented team. Our collaborative culture fosters creativity and ownership, making it a truly rewarding place to build your career.