At a Glance
- Tasks: Shape the future of AI by designing and building HPC systems software.
- Company: Join Nscale, a leading GPU cloud provider for AI innovation.
- Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
- Other info: Diverse and inclusive workplace with a focus on innovation and accountability.
- Why this job: Make a real impact on cutting-edge technology in a collaborative environment.
- Qualifications: Experience in HPC systems and strong programming skills in Go or Python.
The predicted salary is between 80100 - 97900 £ per year.
About Nscale
Nscale is the GPU cloud engineered for AI.
We provide cost-effective, high-performanceinfrastructure for AI start-ups and large enterprise customers.
Nscale enables AI-focusedcompanies to achieve superior results by reducing the complexity of AI development.
Our GPUcloud bolsters technical capabilities and directly supports strategic business outcomes, includingcost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every teammember takes pride in their work and drives it with excellence and urgency.
As an Nscaler, you’llbuild trust through openness and transparency, where everyone is inspired to do their best work.
Ifyou join our team, you’ll be contributing to building the technology that powers the future.
About The Role
We’re hiring a Staff HPC Systems Software Engineer to definethe technical direction and evolution of a core HPC platform domain at Nscale.
In this role, you will operate beyond a single team, shaping how multiple teams build, automate, and run Slurm-based capabilities within Nscale’s wider cloud-native platform.
You’ll work acrossengineering boundaries to bring coherence to architecture, interfaces, lifecycle models, andoperational approaches, while partnering closely with teams working on platform tooling, infrastructure APIs, identity systems, and Kubernetes-adjacent systems.
This is a high-impact staff-level role for someone who combines deep hands-on softwareengineering with strong systems judgement.
Your work will help ensure Nscale’s HPC services arerobust, supportable, and maintainable, while creating leverage through shared patterns, reusableimplementations, and clear technical direction across ambiguous, business-critical problemspaces.
- What You'll Be Doing
- Domain Architecture & Technical Direction
- Own and evolve the technical direction for a defined HPC systems domain, such as Slurmplatform architecture, scheduler integrations, cluster lifecycle, workload environments orservice automation.
- Make architectural decisions that balance software quality, operational realities, customerneeds, and long-term maintainability.
- Define how proven Slurm implementations should be packaged, automated and exposedas a service.
- Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operatingmodels across teams.
- Act as the technical escalation point for the most complex issues within the domain.
- Cross-Team Engineering Leverage
- Establish shared patterns for automation, service lifecycle management, observability, reliability and supportability across the HPC platform.
- Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems and platform tooling.
- Create reusable modules, automation, deployment patterns, and referenceimplementations that increase engineering leverage.
- Identity and correct avoidable technical divergence, duplicated effort and fragile operatingmodels
- Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation and production operations.
- Delivery, Reliability & Influence
- Lead technically critical initiatives spanning 2-4 teams or a defined HPC platform area.
- Unblocked delivery by clarifying technical direction and reducing ambiguity in complexsystem design problems.
- Contribute hands-on where needed to de-risk or accelerate critical work.
- Influence engineering teams without formal authority through strong judgement, designclarity and practical solutions.
- Partner with adjacent cloud-native software engineers so HPC implementations build onshared platform patterns rather than separate ones.
- KPIs
- Technical direction across a defined HPC domain
- Delivery of critical initiatives across 2-4 teams
- Reduction in technical divergence and duplicated effort
- Reliability and supportability of Slurm-based HPC services
About You
- Extensive experience designing and building production software and automation for HPCsystems, especially Slurm-based environments.
- Strong track record of writing maintainable, testable and resilient software in Go, Python orsimilar languages.
- Proven ability to define technical direction across a domain spanning multiple teams orservices.
- Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concernsand operational trade offs.
- Strong practical understanding of GPU-backed infrastructure and HPC networking, including Infini Band, Ro CE, RDMA and performance sensitive workload characteristics.
- Experience integrating HPC systems with cloud-native platforms, APIs, or service deliverymodels.
- Experience creating engineering leverage through standards, reusable patterns, sharedtooling and architectural clarity.
- Strong judgement in balancing short-term delivery with long-term platform health andsupportability.
- Strong written and verbal communication skills, with the ability to align multiple teamsaround a coherent technical direction.
- Experience with other schedulers or batch systems such as Kueue is valuable.
- What We Can Offer You
You’ll have the opportunity to help shape the operating standards behind a next-generation AI
cloud platform, working on complex infrastructure challenges with real ownership and impact.
This
is a chance to play a meaningful role in scaling high-performance, sustainable data centre
operations in a fast-moving environment
Equal Opportunities Statement
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position.
The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
#J-18808-Ljbffr
Staff HPC Systems Software Engineer employer: AI Startups UK
Lovable is an exceptional employer for data scientists, particularly in our dynamic London growth team. With a culture of extreme ownership and low-ego collaboration, we empower our employees to drive impactful growth metrics while working alongside talented engineers. Our commitment to innovation and rapid experimentation offers unparalleled opportunities for professional development, making Lovable a truly rewarding place to build your career.
StudySmarter Expert Advice🤫
We think this is how you could land Staff HPC Systems Software Engineer
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at AI Startups UK or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to AI Startups UK.
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like AI Startups UK.
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like AI Startups UK that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Staff HPC Systems Software Engineer
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at AI Startups UK.
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at AI Startups UK and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at AI Startups UK
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If AI Startups UK uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.