Infrastructure / DevOps Lead

Infrastructure / DevOps Lead

Full-Time 80000 - 100000 £ / year (est.) No working from home possible
B

At a Glance

  • Tasks: Lead a team to optimise AI infrastructure and manage cutting-edge GPU resources.
  • Company: Join BRAHMA AI, a leader in generative media technology.
  • Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
  • Other info: Dynamic startup environment with a focus on creativity and collaboration.
  • Why this job: Shape the future of AI while working with innovative technologies and talented teams.
  • Qualifications: Experience in leading DevOps teams and managing GPU-intensive workloads.

The predicted salary is between 80000 - 100000 £ per year.

Description

Brahma AI operates at the intersection of enterprise Media Asset Management (MAM) and cutting-edge generative media.

We build and scale industry-leading generative AI models, including hyper-realistic digital humans (ATMAN) and multilingual voice synthesis (VAANI), for world-class enterprise clients in entertainment, sports, healthcare, and retail.

Role Overview

We are looking for a Lead Infrastructure / Dev Ops Engineer to lead our core AI Platform & Infrastructure team.

Reporting directly to the VP of Engineering, you will guide a team of 7 engineers responsible for powering our high-performance GPU infrastructure, multi-cloud setup (with GCP as our primary provider), ML model training pipelines, and containerised orchestration environments.

This role balances technical leadership, team management, cloud resource optimisation, and high-level architectural oversight.

You will ensure our research and engineering teams have the fast, scalable, and reliable compute environments necessary to train and deploy state-of-the-art AI models.

Key Responsibilities

  • 1. People, Team & Process Leadership (50%)
  • Team Management: Lead, mentor, and grow a team of 7 Dev Ops and Infrastructure engineers through regular 1:1s, performance reviews, and career pathing.
  • Sprint & Operational Delivery: Drive agile delivery, sprint planning, and backlog prioritisation to align infrastructure deliverables with AI research and product roadmaps.
  • Engineering Standards: Establish best practices for Reliability Engineering, Infrastructure-as-Code (Ia C), continuous integration, and incident post-mortems.
  • Cross-Functional Alignment: Act as the primary technical bridge between infrastructure, ML researchers, AI application developers, and the VP of Engineering.
  • 2. Resource Management & Fin Ops (25%)
  • GPU & Multi-Cloud

Management: Manage high-density GPU clusters across a multi-cloud ecosystem (primarily GCP) optimised for large custom AI model training and real-time inference workflows.

  • Fin Ops & Cost Control: Oversee infrastructure consumption, track cloud/hardware costs, negotiate vendor terms, and optimise GPU utilisation to maintain cost efficiency.
  • High-Performance Storage: Oversee high-throughput storage and caching solutions engineered for ultra-fast data retrieval and low-latency access.
  • 3. Architecture, Engineering & Compliance (25%)
  • Technical Escalation & Hands-On Oversight: Serve as the senior technical escalation point for complex infrastructure incidents and architecture decisions.
  • Automation & Ia C: Standardise platform deployments using Infrastructure as Code (e. g., Terraform/Open Tofu) and modern container orchestration (Kubernetes).
  • Security & Compliance

Collaboration: Partner with security stakeholders to ensure our AI training environments meet industry security standards (e. g., MPA Best Practices, ISO 27001, SOC 2).

  • Must Haves
  • Leadership Experience: Proven track record leading or managing a team of 5+ infrastructure, platform, or Dev Ops engineers.
  • AI/GPU Infrastructure: Hands-on experience architecting and managing GPU-intensive workloads (NVIDIA clusters, cloud AI accelerators) for compute-heavy applications.
  • Multi-Cloud & Orchestration: Expertise with Kubernetes, Docker, Terraform (or Open Tofu), and multi-cloud environments (with strong hands-on GCP experience).
  • High-Performance

Storage: Demonstrated experience designing, optimising, and maintaining high-performance storage architectures and caching layers for demanding compute workloads.

  • Cloud & Resource Management: Strong experience with cloud cost governance (Fin Ops), capacity planning, and vendor interaction.
  • Communication: Exceptional stakeholder management skills with the ability to bridge business requirements and deep technical infrastructure details.
  • Nice to Have
  • Experience managing physical data centres, co-location facilities, or hybrid infrastructure environments.
  • Working knowledge of ML orchestration frameworks (e. g., Ray, Slurm, Kubeflow).
  • Background in media pipelines, VFX tooling, or media compliance standards (MPA, ISO 27001).
  • Prior experience working in a hybrid startup/scale-up environment.

About BRAHMA AI

BRAHMA AI is the next generation of enterprise media technology formed through the integration of Prime Focus Technologies and Metaphysic.

By combining CLEAR®, CLEAR® AI, ATMAN, and VAANI into one ecosystem, BRAHMA AI enables enterprises to manage, create, and distribute content with intelligence, security, and efficiency.

Proven, scalable, and enterprise-tested, BRAHMA AI is helping global organizations accelerate growth, efficiency, and creative impact in the AI-powered era.

Infrastructure / DevOps Lead employer: BRAHMA

Brahma AI is an exceptional employer that fosters a dynamic and innovative work culture, where employees are empowered to lead and grow within their roles. With a strong focus on cutting-edge technology and collaboration, the company offers ample opportunities for professional development, particularly in the rapidly evolving fields of AI and media technology. Located in a vibrant tech hub, Brahma AI provides a unique environment that encourages creativity and teamwork, making it an ideal place for those seeking meaningful and rewarding employment.

B

Contact Details:

BRAHMA Recruitment Team

We think you need these skills to ace Infrastructure / DevOps Lead

Team Management
Agile Delivery
Infrastructure-as-Code (IaC)
Continuous Integration
GPU Management
Multi-Cloud Environments
Kubernetes