Referment is working with a cloud infrastructure provider that supports enterprise AI, high-performance computing and cloud-native workloads. The business is looking for a senior technical specialist to shape large-scale GPU platforms for customers across the UK.
The Role
- Design enterprise GPU clusters and produce clear high-level and low-level designs, rack and network diagrams, and detailed bills of materials.
- Shape scale-up and scale-out architectures across NVIDIA GPU platforms, NVLink and NVSwitch, high-speed InfiniBand and RoCE fabrics, storage, power and cooling.
- Translate customer workload, performance and commercial requirements into practical bare-metal, Slurm or Kubernetes-based solutions.
- Lead technical discovery sessions, architecture workshops and executive presentations, and contribute to complex proposals and RFP responses.
- Plan and oversee proofs of concept, using appropriate benchmarking to validate throughput, latency and distributed training performance.
- Work closely with enterprise customers, hardware vendors and internal engineering teams, feeding technical insight back into the platform roadmap.
What Weβre Looking For
- At least five years in solution architecture, systems engineering or technical pre-sales focused on HPC, AI infrastructure or high-performance cloud platforms.
- Deep knowledge of NVIDIA HGX or DGX systems, GPU interconnects and modern rack-scale GPU architectures.
- Strong experience designing InfiniBand and/or RoCE networks, including congestion management and GPU-direct technologies.
- Practical understanding of GPU workload orchestration through Kubernetes and associated NVIDIA operators, or bare-metal environments using technologies such as Slurm, Ansible and Terraform.
- A track record of producing robust HLDs, LLDs, network diagrams and itemised infrastructure designs.
- Confident communication with CTOs, infrastructure leaders and ML engineers, with the judgement to explain hardware, networking and cost trade-offs.
- Awareness of high-density data-centre power, cooling and storage considerations.
Relevant Desirable Experience
NVIDIA AI infrastructure, networking or InfiniBand certifications would be useful, as would experience with performance testing for distributed AI workloads. A relevant degree is welcome, although equivalent practical experience is equally valuable.
This could suit a Senior Solutions Architect, HPC Systems Engineer or technical pre-sales specialist who has designed GPU clusters and wants to remain close to both customers and engineering. You must be based in the UK.
#Referment
#J-18808-Ljbffr
Senior Solution Engineer β GPU and AI Infrastructure (F673F3F) employer: Referment
Referment is an exceptional employer, offering a dynamic work culture that fosters collaboration and innovation. As a Finance Director, you will not only lead the finance function but also play a pivotal role in shaping the company's growth trajectory, with ample opportunities for professional development and advancement. Located in a vibrant B2B technology hub, the company provides a supportive environment where your contributions are valued and rewarded.