Technical Site Reliability Engineer in London

Technical Site Reliability Engineer in London

London Full-Time 60000 - 80000 £ / year (est.) No working from home possible
Anduril Industries, Inc.

At a Glance

  • Tasks: Design and maintain simulation infrastructure for military wargaming facilities.
  • Company: Anduril Industries, a cutting-edge defence technology company.
  • Benefits: Competitive salary, equity grants, and comprehensive health benefits.
  • Other info: Dynamic role with opportunities for growth in a multinational team.
  • Why this job: Join a mission-driven team transforming military capabilities with advanced technology.
  • Qualifications: Proficiency in Python, networking fundamentals, and experience in production systems.

The predicted salary is between 60000 - 80000 £ per year.

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting‑edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

ABOUT THE TEAM

Advanced Capabilities is the internal warfighter research team for Anduril's Maneuver Dominance division. We invent new products, inform vehicle specifications, define autonomous tactics, and work out how humans and teams of autonomous systems will operate together in future contested multi-domain environments. You'll join a small, multinational team of engineers spanning multiple disciplines such as wargaming, game engineering, HPC simulations, LLM agents and VR environments; all to give our warfighters and researchers the leverage to explore faster, test more ideas, and better understand tomorrow’s war.

ABOUT THE ROLE

We are building the next generation of wargaming facilities purpose-built for the military to run massive‑scale simulations of autonomous systems operating in contested environments. Those facilities are only useful if they are up, current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers’ real time.

As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation software and distributed systems. You will be the person who knows why the simulation broke, who catches it before anyone else notices, and who makes sure it doesn’t break the same way twice. This is a hands‑on role that blends software maintenance, infrastructure ownership, and release validation, with direct exposure to the development teams whose code you're keeping stable.

This position is in Abu Dhabi, UAE, with initial position hiring occurring in London, UK. Candidate must be willing to relocate to the facility upon completion.

WHAT YOU'LL DO

  • Maintain the simulation software stack - installation, configuration, updates, version management, and day-to-day functionality across the Simulation Center's tools and environments.
  • Own the underlying infrastructure - compute, networking, storage, and environment configuration that the simulation depends on; keep it provisioned, patched, and performant.
  • Build and maintain a post-release test suite - design, automate, and continually extend a regression and smoke-test process that runs after every software release or configuration change, so integration issues surface immediately rather than mid-exercise.
  • Forecast, diagnose, and eliminate failure modes - root‑cause errors and bugs in the system, drive them to permanent resolution, and implement the guardrails, monitoring, or process changes that prevent recurrence.
  • Partner with development teams and stakeholders - review upcoming changes for reliability risk, surface concerns early, and implement mitigation strategies before releases land in the simulation environment.
  • Monitor overall system health - instrument and watch the environment, triage issues within your scope, and clearly and quickly with the context needed for others to act when an issue exceeds your ability to resolve it.
  • Document what you learn - runbooks, known issues, environment configuration, and release validation results, so the Simulation Center's operational knowledge isn't held in one person's head.

REQUIRED QUALIFICATIONS

  • Proficiency in Python for automation, tooling, and test development.
  • Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
  • Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.
  • Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
  • Demonstrated experience maintaining production or production‑adjacent systems, including troubleshooting under time pressure.
  • Strong written and verbal communication; you can elevate an issue, explain a root cause, and write a runbook someone else can follow.
  • Eligibility to pass the security and background check requirements for sensitive information systems.

PREFERRED QUALIFICATIONS

  • Experience with modeling and simulation, wargaming, or distributed simulation standards (DIS, HLA, TENA) and platforms such as AFSIM, VBS, or similar.
  • Test automation and CI/CD experience; building automated validation pipelines, not just running them.
  • Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).
  • On‑prem and cloud deployment experience.
  • Observability tooling: Prometheus, Grafana, ELK, or equivalent.
  • Prior work in a defense, aerospace, or classified environment.
  • Active security clearance.

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top‑tier benefits for full‑time employees, including: At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next.

Technical Site Reliability Engineer in London employer: Anduril Industries, Inc.

Anduril Industries is an exceptional employer, offering a dynamic work environment in Llanbedr, Wales, where innovation meets defence technology. Employees benefit from competitive salaries, comprehensive health coverage, and professional development opportunities, all while contributing to groundbreaking projects that shape the future of military capabilities. The collaborative culture fosters growth and encourages team members to push the boundaries of technology, making it a rewarding place for those passionate about advancing autonomous systems.

Anduril Industries, Inc.

Contact Details:

Anduril Industries, Inc. Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Technical Site Reliability Engineer in London

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Anduril Industries, Inc. or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Anduril Industries, Inc..

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Anduril Industries, Inc..

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Anduril Industries, Inc. that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Technical Site Reliability Engineer in London

Python
C++
Networking Fundamentals
TCP/IP
UDP
DNS
Routing

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Anduril Industries, Inc..

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Anduril Industries, Inc. and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at Anduril Industries, Inc.

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Anduril Industries, Inc. uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.