Software Engineer, GPU Infrastructure- ChatGPT Engineering

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Full-Time 80000 - 100000 £ / year (est.) No working from home possible
Doist

At a Glance

  • Tasks: Design and build software for managing large-scale GPU infrastructure supporting ChatGPT.
  • Company: Join a leading AI company powering one of the world's largest AI products.
  • Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
  • Other info: Collaborative environment with a focus on innovation and operational excellence.
  • Why this job: Make an impact on cutting-edge AI technology while solving complex operational challenges.
  • Qualifications: 5+ years in software engineering with experience in GPU or compute infrastructure.

The predicted salary is between 80000 - 100000 £ per year.

About the

Team: Chat GPT Engineering builds and operates the compute platform powering one of the world's largest AI products.

Every Chat GPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance.

As our GPU fleet continues to grow, we’re investing in the infrastructure that operates it.

Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous.

We work across production engineering, distributed systems, capacity management, and AI‑powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.

About the Role

We’re looking for a Software Engineer with deep experience operating large‑scale GPU or compute infrastructure.

You’ll design and build the systems that manage GPU clusters at scale—from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention.

You’ll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization.

This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI.

  • In This Role, You Will
  • Design, build, and operate software that manages large‑scale GPU infrastructure supporting Chat GPT inference.
  • Build internal platforms, tooling, and AI‑powered agents that automate fleet operations and reduce operational overhead.
  • Improve observability, reliability, and operational efficiency across thousands of GPUs.
  • Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
  • Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.
  • Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform.
  • Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.
  • You Might Thrive in This Role
  • Have experience operating large‑scale production infrastructure, preferably GPU clusters or other compute‑intensive distributed systems.
  • Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering.
  • Have built software that automates operational workflows rather than relying on manual processes.
  • Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure.
  • Understand infrastructure observability, monitoring, capacity planning, and incident management.
  • Enjoy identifying cross‑team pain points and building reusable platforms that improve developer productivity.
  • Are comfortable working across software engineering and systems operations, owning problems end‑to‑end.
  • Thrive in fast‑moving environments with significant technical ambiguity.

Qualifications

  • 5+ years of software engineering experience building production infrastructure.
  • Strong programming skills in Go, Python, C++, Rust, or similar systems languages.
  • Experience designing and operating highly available distributed systems.
  • Experience with GPU infrastructure, high‑performance computing, ML infrastructure, or large‑scale compute platforms.
  • Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
  • Excellent debugging, systems design, and operational problem‑solving skills.
  • Strong communication skills and experience collaborating across engineering organizations.

We are an equal‑opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

#J-18808-Ljbffr

Software Engineer, GPU Infrastructure- ChatGPT Engineering employer: Doist

Brink's is an exceptional employer that fosters a collaborative and innovative work culture, empowering employees to drive global commercial strategies in the ATM lifecycle solutions sector. With a strong focus on professional development and growth opportunities, employees are encouraged to expand their skills while contributing to sustainable growth and operational excellence. Located in a dynamic environment, Brink's offers unique advantages such as a diverse team and the chance to build long-term partnerships with global customers, making it a rewarding place to advance your career.

Doist

Contact Details:

Doist Recruitment Team

We think you need these skills to ace Software Engineer, GPU Infrastructure- ChatGPT Engineering

GPU Infrastructure Management
Large-Scale Production Infrastructure
Operational Automation
Capacity Planning
Incident Response
Kubernetes
Linux Systems