Infrastructure engineer (UK) in London

Infrastructure engineer (UK) in London

London Full-Time 56700 - 69300 £ / year (est.) Home office (partial)
United States Digital Space LLC

At a Glance

  • Tasks: Design and maintain resilient infrastructure for AI-powered workflows.
  • Company: Join a leading enterprise in generative AI with a collaborative culture.
  • Benefits: Generous PTO, medical insurance, parental leave, and wellness stipends.
  • Other info: Dynamic hybrid work environment with excellent career growth opportunities.
  • Why this job: Be at the forefront of AI innovation and make a real impact.
  • Qualifications: 5+ years in infrastructure engineering and experience with AI tools.

The predicted salary is between 56700 - 69300 £ per year.

About the company the company is where the world's leading enterprises orchestrate AI-powered work.

Our vision is to expand human capacity through superintelligence.

And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation.

With the company's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by the company's enterprise-grade LLMs.

Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, the company is rapidly cementing its position as the leader in enterprise generative AI.

Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI.

About the role

At the company, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7.

As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows.

This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI.

You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering.

What you'll do

  • Technical
  • Breadth across disciplines.

Bring deep focus to one problem at a time, with the breadth to move between SRE, Dev Ops, Infrastructure, and Platform work over a quarter or two as the leverage shifts.

This is not a thrash-every-week role — most of the time you're heads-down on one substantial initiative (the on-call posture, the release pipeline, the multi-region Terraform layout, the internal platform surface).

Cross-layer fluency is what lets you pick the right next initiative; it isn't a weekly context-switch.

  • Simplicity / via negativa.

Challenge the status quo and remove toil before adding features — automate operational tasks and infrastructure management with Python or Go, reject tools that don't fit the problem, and treat manual on-call work as a defect to be designed out, not a status quo to be staffed up.

  • Breadth across the stack.

Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and the supporting cloud and AI tooling that backs the company's high-traffic platform.

  • AI in workflow.

Run agents in your daily loop — Claude Code, Droid, Codex, internal skills — to investigate incidents, draft Terraform / Helm changes, write runbooks, scaffold tooling, and review PRs.

Build the agentic setup as a collective surface: humans and digital teammates working as one team, with shared skills, shared context, and shared on-call workflows.

Encode recurring infra tasks as internal skills any teammate (human or agent) can pick up and run, so the team's throughput compounds — not just your own.

  • Debugging fluency.

Lead incident response, post-mortems, and root-cause analyses — trace failures to the underlying problem (never the symptom), apply the learning back into the architecture, and prevent the same incident from happening twice.

  • Non-technical
  • End-to-end ownership.

Own the reliability, performance, and efficiency of the company's core services end-to-end — define and uphold the SLOs and error budgets, carry the on-call pager, and stand behind the outcome metric, not just the system you shipped.

  • Strategic vs. tactical balance.

Balance this week's critical work with the 6–12-month platform direction — ship the on-call-driving fix today while shaping the multi-year observability, cost, and reliability investments that move the company's enterprise customers.

  • Cross-functional collaboration.

Operate at the seams with product, security, and engineering peers — provide expert guidance on system design for reliability, performance, and scalability from conception through launch, Connect the infra agenda to product and revenue context, and disagree with evidence, not volume.

  • What you need
  • Technical
  • Track record. 5+ years of experience in infrastructure engineering, Dev Ops, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company.
  • Breadth.

Experience running containerisation in production (a real cluster, not a lab), with experience in Helm and Terraform or Pulumi on at least one major cloud (AWS preferred), plus good proficiency in Python or Go for automation and tooling.

  • AI in workflow.

AI is part of how you ship, not a thing you've read about — agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop, you've built or adopted AI-assisted workflows others now use, and you have strong opinions on where it's unreliable.

This is a hard requirement, not a bonus.

Candidates whose actual daily workflow does not already include AI tooling will not be advanced.

  • First-principles + decision-making.

Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems — reason from constraints and failure modes (not analogy or vendor defaults), name the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off), and reject the "best practices" answer when it doesn't fit the problem.

  • Reversibility & blast-radius.

Make reversible calls by default — write the rollback before you touch production, work fluently with monitoring and logging stacks (Prometheus, Grafana, ELK or equivalent), and stress the system in safe places so it comes back stronger.

  • Non-technical
  • Cross-functional collaboration.

Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams — surface non-goals before anyone asks, and partner with product, security, and platform peers as one delivery surface.

  • Autonomy & end-to-end ownership.

A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability.

At least one 0-to-1 infrastructure build you owned end-to-end, with the outcome metric attached.

  • Bonus if you have
  • Software-engineering depth.

A software-engineering background, not only config and scripting — you've designed, built, and shipped non-trivial production code (services, libraries, internal frameworks) in Python, Go, or a comparable language, you can read and modify the codebases your infrastructure runs, and you move between infra automation and feature engineering without changing brains.

Benefits & perks (UK full-time employees)

  • Generous PTO, plus company holidays
  • Comprehensive medical and dental insurance
  • Paid parental leave for all parents (16 weeks)
  • Fertility and family planning support
  • Early-detection cancer testing through Galleri
  • Competitive pension scheme and company contribution

• Annual work-life stipends for

  • Wellness stipend for gym, massage/chiropractor, personal training, etc.
  • Learning and development stipend
  • Company-wide off-sites and team off-sites
  • Competitive compensation and company stock options
  • #J-18808-Ljbffr

Infrastructure engineer (UK) in London employer: United States Digital Space LLC

United States Digital Space LLC is an exceptional employer, offering a dynamic work culture that prioritises innovation and collaboration in the heart of Greater London. With a strong focus on employee well-being and flexible work options, we provide ample opportunities for professional growth and development, making it an ideal environment for those looking to make a meaningful impact in the field of AI-enabled SaaS engineering.

United States Digital Space LLC

Contact Details:

United States Digital Space LLC Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Infrastructure engineer (UK) in London

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at United States Digital Space LLC or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to United States Digital Space LLC.

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like United States Digital Space LLC.

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like United States Digital Space LLC that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Infrastructure engineer (UK) in London

Infrastructure Engineering
DevOps
Containerisation
AWS
GCP
Azure
Kubernetes

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at United States Digital Space LLC.

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at United States Digital Space LLC and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at United States Digital Space LLC

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If United States Digital Space LLC uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.