Principal Infrastructure Engineer, AI Cluster Performance & Validation in London

Principal Infrastructure Engineer, AI Cluster Performance & Validation in London

London Full-Time 90000 - 110000 £ / year (est.) No working from home possible
U

At a Glance

  • Tasks: Lead the performance and validation of cutting-edge AI clusters in a dynamic tech environment.
  • Company: Join a leading AI infrastructure team focused on innovation and excellence.
  • Benefits: Enjoy competitive pay, flexible work options, and opportunities for professional growth.
  • Other info: Collaborative culture with mentorship opportunities and a focus on operational excellence.
  • Why this job: Make a real impact on AI technology while working with the latest advancements.
  • Qualifications: 10+ years in large-scale compute infrastructure and hands-on AI workload experience required.

The predicted salary is between 90000 - 110000 £ per year.

As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs.

Key Responsibilities

  • Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover.
  • Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Run and instrument real AI workloads as a diagnostic instrument.
  • Stand up and execute distributed training and inference jobs - open-source and customer-representative models - across thousands of accelerators to validate cluster behavior under genuine load rather than synthetic proxies alone, and translate what those runs reveal into fleet-wide fixes.
  • Lead deep diagnosis of large-scale cluster failures and performance regressions, isolating root cause across the full stack: GPU and NIC firmware, PCIe/NVLink topology and NUMA placement, InfiniBand/RoCE fabric health, congestion control and routing, storage and data-loader throughput, scheduler placement, and framework/communication-library behavior.
  • Serve as the final escalation point for the hardest slow-job and stalled-job investigations.
  • Design and build the validation and burn-in systems that qualify nodes, racks, and full pods at scale — NCCL/RCCL collective sweeps, HPL/HPCG and MLPerf-style benchmarks, thermal and power soak tests, straggler and flapping-link detection — and automate them so that qualification is a repeatable pipeline, not a manual campaign.
  • Drive cluster optimization end to end, tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput.
  • Partner with Infrastructure, Platform, SRE, and customer-facing teams to translate operational and customer performance needs into durable engineering solutions, and to feed diagnostic signal back into provisioning, remediation, and capacity workflows.
  • Build production-grade Python systems and performance tooling for automated triage, telemetry correlation, and regression detection, leveraging AI tools to accelerate delivery.
  • Assess impact to the team's software and validation stack from new hardware product programs, and explore AI-driven process improvement and automation.
  • Establish engineering standards for reliability, observability, benchmarking methodology, and operational excellence across all services, and raise the diagnostic capability of the wider organization through mentorship, runbooks, and post-incident technical write-ups.

Required Qualifications

  • Education: Bachelor's or higher degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Experience: 10+ years of relevant experience building, operating, or debugging large-scale compute infrastructure, including significant time at staff or principal level owning cross-team technical direction.
  • AI Workload Expertise: Hands-on experience running real AI compute jobs at scale - pre-training, fine-tuning, or large-scale inference of open-source or proprietary models - including practical familiarity with distributed training strategies (data, tensor, pipeline, and expert parallelism) and frameworks such as PyTorch, Megatron-LM, DeepSpeed, or equivalent.
  • Cluster Validation: Demonstrated experience validating and accepting large clusters (thousands of GPUs) for performance and reliability, with a working command of benchmark methodology and the ability to defend a number to both engineers and customers.
  • Performance Debugging: Proven ability to diagnose distributed performance problems - stragglers, collective stalls, link flaps, thermal throttling, silent data corruption, ECC and Xid errors, noisy-neighbour and storage-bound bottlenecks - using tools such as NCCL debug tracing, Nsight Systems/Compute, PyTorch Profiler, perf, and fabric telemetry.
  • Networking: Deep understanding of high-performance fabrics - InfiniBand and/or RoCEv2, RDMA, GPUDirect, adaptive routing, congestion control, and rail-optimized topologies - and of networking fundamentals (TCP/IP, BGP).
  • Systems & Programming: Deep Linux systems expertise (kernel tunables, NUMA, PCIe, IRQ and memory behaviour) and strong production Python, plus experience with C/C++ or Go and with configuration management tooling (e.g., Ansible, Terraform).
  • Schedulers: Experience operating and debugging AI workloads under SLURM and/or Kubernetes at scale.

Preferred Qualifications

  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience bringing up and qualifying a greenfield GPU supercluster from first rack to production traffic, including firmware, driver, and topology standardization across a heterogeneous fleet.
  • Direct experience with NVIDIA GPU platforms (H200/GB200/GB300-class), NVLink and NVSwitch domains, DCGM, SHARP, UFM, and the NVIDIA software stack; or equivalent depth on AMD Instinct and ROCm/RCCL.
  • Published or presented benchmark, scaling, or post-mortem work - MLPerf submissions, scaling studies, or public technical write-ups on large-cluster behaviour.
  • Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) applied to high-cardinality GPU and fabric telemetry, including automated anomaly and regression detection.
  • Experience with high-throughput parallel storage (Lustre, GPFS, WEKA, VAST) and with diagnosing data-pipeline-bound training jobs.
  • Familiarity with cloud-native technologies (Kubernetes, Docker), infrastructure-as-code principles, and integration with infrastructure tooling such as DCIMs, NetBox, and bare metal APIs (MAAS, Ironic, IPMI, Redfish).
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
  • Familiarity with SLOs/metrics measurement and logs/telemetry/metrics integration with tools for enhanced operator experience.

Principal Infrastructure Engineer, AI Cluster Performance & Validation in London employer: UNCOVER

Join a prestigious international law firm in London, where you will be part of a high-performing legal team dedicated to excellence and innovation. The firm offers a collaborative work culture that fosters professional growth, providing ample opportunities for career advancement and skill development. With a focus on supplier agreements and commercial technology contracts, this role not only promises meaningful work but also the chance to make a significant impact within a globally recognised organisation.

U

Contact Details:

UNCOVER Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Principal Infrastructure Engineer, AI Cluster Performance & Validation in London

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at UNCOVER or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to UNCOVER.

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like UNCOVER.

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like UNCOVER that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Principal Infrastructure Engineer, AI Cluster Performance & Validation in London

AI Workload Expertise
Distributed Training Strategies
Benchmark Methodology
Performance Debugging
High-Performance Fabrics Knowledge
Linux Systems Expertise
Production Python Programming

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at UNCOVER.

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at UNCOVER and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at UNCOVER

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If UNCOVER uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.