GenAI Infra Engineer: Real-Time GPU Serving & Open-Weights

GenAI Infra Engineer: Real-Time GPU Serving & Open-Weights

Full-Time 80000 - 100000 Β£ / year (est.) No working from home possible
Deliveroo

At a Glance

  • Tasks: Build cutting-edge infrastructure for AI model serving and inference pipelines.
  • Company: Join Deliveroo's innovative GenAI Platform team.
  • Benefits: Competitive salary, flexible working, and opportunities for professional growth.
  • Other info: Dynamic team environment focused on innovation and collaboration.
  • Why this job: Make a real impact in AI by developing scalable solutions.
  • Qualifications: 3+ years in software engineering with strong Python and distributed systems skills.

The predicted salary is between 80000 - 100000 Β£ per year.

Deliveroo's GenAI Platform team is hiring a Software Engineer to build production infrastructure for open-weights model serving, inference pipelines, and fine-tuning. You will work across backend services, GPU autoscaling, and observability to deliver scalable, cost-aware AI platforms for internal and partner teams.

Ideal candidates have:

  • 3+ years in software engineering
  • Strong Python and distributed systems skills
  • Hands-on experience with ML infrastructure at scale and production-grade

GenAI Infra Engineer: Real-Time GPU Serving & Open-Weights employer: Deliveroo

Deliveroo is an excellent employer that fosters a vibrant work culture in Leeds, where innovation and collaboration thrive. With a strong focus on employee growth, we offer comprehensive training and development opportunities, ensuring our team members can advance their careers while making a meaningful impact in the local community. Join us to be part of a dynamic marketplace that values your contributions and rewards your efforts.

Deliveroo

Contact Details:

Deliveroo Recruitment Team

We think you need these skills to ace GenAI Infra Engineer: Real-Time GPU Serving & Open-Weights

Software Engineering
Python
Distributed Systems
Machine Learning Infrastructure
Production-Grade Systems
GPU Autoscaling
Observability