At a Glance
- Tasks: Design and maintain simulation infrastructure for military wargaming facilities.
- Company: Join Anduril, a leader in innovative defence technology.
- Benefits: Competitive salary, equity grants, and comprehensive health benefits.
- Other info: Exciting opportunity to work in a dynamic, multinational team.
- Why this job: Make a real impact on military simulations and work with cutting-edge tech.
- Qualifications: Proficient in Python, with networking and troubleshooting skills.
The predicted salary is between 59400 - 72600 £ per year.
ABOUT THE TEAM
Advanced Capabilities is the internal warfighter research team for Anduril's Maneuver Dominance division. We invent new products, inform vehicle specifications, define autonomous tactics, and work out how humans and teams of autonomous systems will operate together in future contested multi-domain environments. You'll join a small, multinational team of engineers spanning multiple disciplines such as wargaming, game engineering, HPC simulations, LLM agents and VR environments; all to give our warfighters and researchers the leverage to explore faster, test more ideas, and better understand tomorrow’s war.
ABOUT THE ROLE
We are building the next generation of wargaming facilities purpose-built for the military to run massive-scale simulations of autonomous systems operating in contested environments. Those facilities are only useful if they are up, current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers’ real time. As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation software and distributed systems. You will be the person who knows why the simulation broke, who catches it before anyone else notices, and who makes sure it doesn't break the same way twice. This is a hands-on role that blends software maintenance, infrastructure ownership, and release validation, with direct exposure to the development teams whose code you're keeping stable. This position is in Abu Dhabi, UAE, with initial position hiring occurring in London, UK. Candidate must be willing to relocate to the facility upon completion.
WHAT YOU'LL DO
- Maintain the simulation software stack - installation, configuration, updates, version management, and day-to-day functionality across the Simulation Center's tools and environments.
- Own the underlying infrastructure - compute, networking, storage, and environment configuration that the simulation depends on; keep it provisioned, patched, and performant.
- Build and maintain a post-release test suite - design, automate, and continually extend a regression and smoke-test process that runs after every software release or configuration change, so integration issues surface immediately rather than mid-exercise.
- Forecast, diagnose, and eliminate failure modes - root-cause errors and bugs in the system, drive them to permanent resolution, and implement the guardrails, monitoring, or process changes that prevent recurrence.
- Partner with development teams and stakeholders - review upcoming changes for reliability risk, surface concerns early, and implement mitigation strategies before releases land in the simulation environment.
- Monitor overall system health - instrument and watch the environment, triage issues within your scope, and escalate clearly and quickly with the context needed for others to act when an issue exceeds your ability to resolve it.
- Document what you learn - runbooks, known issues, environment configuration, and release validation results, so the Simulation Center's operational knowledge isn't held in one person's head.
REQUIRED QUALIFICATIONS
- Proficiency in Python for automation, tooling, and test development.
- Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
- Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.
- Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
- Demonstrated experience maintaining production or production-adjacent systems, including troubleshooting under time pressure.
- Strong written and verbal communication; you can escalate an issue, explain a root cause, and write a runbook someone else can follow.
- Eligibility to pass the security and background check requirements for sensitive information systems.
PREFERRED QUALIFICATIONS
- Experience with modeling and simulation, wargaming, or distributed simulation standards (DIS, HLA, TENA) and platforms such as AFSIM, VBS, or similar.
- Test automation and CI/CD experience; building automated validation pipelines, not just running them.
- Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).
- On-prem and cloud deployment experience.
- Observability tooling: Prometheus, Grafana, ELK, or equivalent.
- Linux systems administration depth; comfort in mixed Linux/Windows environments.
- Prior work in a defense, aerospace, or classified environment.
- Active security clearance.
The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full-time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees.
At Anduril, we invest in our people. Our comprehensive, competitive benefits package ensures you’re supported in health, recovery, and whatever comes next.
Technical Site Reliability Engineer in London employer: Anduril
Anduril is an exceptional employer that fosters a dynamic work culture in Greater London, where innovation meets purpose. Employees benefit from a collaborative environment that encourages professional growth through mentorship and leadership opportunities, particularly for those passionate about advancing defense technology. With a commitment to impactful projects and cutting-edge autonomy solutions, Anduril offers a unique chance to contribute to critical systems while being part of a forward-thinking team.
StudySmarter Expert Advice🤫
We think this is how you could land Technical Site Reliability Engineer in London
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Anduril or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Anduril.
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Anduril.
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Anduril that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Technical Site Reliability Engineer in London
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Anduril.
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Anduril and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at Anduril
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Anduril uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.