Nebius is seeking an experienced Site Reliability Engineer to own the reliability, performance, and observability of the full inference stack. You will design telemetry pipelines, tune Kubernetes autoscalers, and craft Terraform modules to ensure cost efficiency and resilience.
You’ll respond to incidents, drive post-morts, and collaborate with software engineers to turn reliability into a product feature. This role requires deep experience with Kubernetes, Prometheus, Grafana, Terraform, and
#J-18808-Ljbffr
Senior SRE: AI Cloud Reliability & GPU Scale employer: Nebius Group
Nebius is an exceptional employer, offering a dynamic and innovative work environment in the heart of London. With a strong focus on employee wellbeing and professional growth, we provide competitive compensation, flexible working arrangements, and the opportunity to contribute to groundbreaking AI projects alongside a talented team. Our collaborative culture fosters creativity and ownership, making it a truly rewarding place to build your career.