At a Glance
- Tasks: Champion reliability engineering and improve AI network quality through innovative processes and data metrics.
- Company: Fluidstack, a pioneering tech company focused on expanding human freedom through AI.
- Benefits: Competitive salary, equity options, health insurance, and generous PTO.
- Other info: Collaborative environment with opportunities for growth and travel.
- Why this job: Join us in building civilization-scale infrastructure for AI and make a real impact.
- Qualifications: 5+ years in network engineering with strong operational and software development experience.
The predicted salary is between 120000 - 200000 £ per year.
About Fluidstack
We exist to make humanity more free. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.
About The Role
Fluidstack is seeking a Network Engineer, Reliability & Observability to serve as a reliability engineer championing and building process, data collections, and reliability metrics with the objective of improving the quality and reliability of AI networks from deployment through the full lifecycle of operations. This role is focused on developing processes, systems, tools, data and data pipelines, and observability to improve the quality of networks and deliver automated metrics (24x7) as well as periodic reliability reports for both internal and external customers.
This role is ideal for experienced network operators who are passionate about reliability and have experience designing and building full lifecycle software such as Quality Assurance audits, circuit audits, periodic audits, failure rates and failure analysis. You are passionate about hardware (electronics and optics), software development, and you value and promote the use of data to make informed decisions in deployment, operations, and strategic sourcing. Experienced SRE (Site Reliability Engineers) with a passion for networking are encouraged to apply.
Focus
- Ownership of Quality Assurance: Design, develop, and support QA process for network hardware and networks.
- Pipelines: Develop and deploy serverless workflows, server based, and manually triggered data pipelines producing network quality and reliability observability for internal and external customers.
- Deployment and Operations Support: Support full lifecycle data collection and analysis partnering with Deployment, Operations, DC hardware, and logistics teams to produce data that drives process improvements and delivers on SLA and SLOs.
- Process Engineering: Develop, pilot, and deploy process improvements for deployment and repair to produce data and consume data with Machine Learning to fulfill our mission.
- Cross-Team Collaboration: Own without ego and execute in a collaborative team with design, deployment, operations engineers and software developers.
- Subject Matter Expert: In at least two or more deep subjects such as IP routing, optics, optical transport, Ethernet, RDMA/RoCE, or electrical power.
About You
- Strong Operations Background: 5+ years in network engineering and at least 3+ years in operations with significant hands‑on operational experience. You've run production networks or compute, responded to incidents at all hours, and debugged complex failures under pressure. You understand the difference between "working" and "production-ready".
- Software Development: You have experience with ITIL, Agile (xP), and TDD including developing and leading programs and projects. You have experience building hyperscale platforms, demonstrating a fluency in Golang with supporting tools in Python or RUST.
- Datacenter Fabric Expertise: Deep experience operating modern datacenter networks including EVPN/VXLAN, BGP, CLOS topologies, and high‑radix switching. You’re comfortable troubleshooting Layer 2/3 issues, BGP routing problems, fabric misconfigurations, and physical media failures.
- Incident Response Excellence: Proven ability to lead incident response, perform systematic troubleshooting, and drive issues to resolution. You remain calm during outages, communicate clearly with stakeholders, and know when to escalate versus when to dig deeper.
- Matrix Leadership Experience: You understand how to build relationships with onsite teams, coordinate physical infrastructure work, and represent network engineering in a field environment.
- Operational Pragmatism: You balance perfection with progress. You can troubleshoot with imperfect information, make pragmatic decisions under time pressure, and prioritize based on business impact.
- Self Driven: You embrace complex challenges with undefined processes and key results. You can dive in to learn, but zoom back out to build Objectives, develop Key Results and build a software development project and pipeline in Jira solo.
- Travel: You are willing and able to travel to spend time with the team at our local offices or data center locations, up to 20% of the time.
Nice to Haves
- AI/HPC Fabric Operations: Experience operating AI/ML or HPC fabrics with RDMA (RoCEv2), lossless Ethernet (PFC, ECN), or high‑performance networking.
- Reliability Engineering: You have experience with observability and reliability engineering from network operations or in manufacturing quality.
- Hardware Repair Experience: Hands‑on experience coordinating hardware repairs, RMAs, and physical infrastructure work.
- Observability & Monitoring: Familiarity with network monitoring platforms, alerting systems, and telemetry collection.
Salary & Benefits
Competitive total compensation package (salary + equity). Retirement or pension plan, in line with local norms. Health, dental, and vision insurance. Generous PTO policy, in line with local norms.
The base salary range for this position is $150,000 - $250,000 per year, depending on experience, skills, qualifications, and location. This range represents our good faith estimate of the compensation for this role at the time of posting. Total compensation may also include equity in the form of stock options.
We are committed to pay equity and transparency. Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law.
Network Engineer, Reliability & Observability in London employer: FluidStack
Fluidstack is an exceptional employer dedicated to empowering its employees through a culture of innovation and collaboration. With a focus on building civilization-scale infrastructure for AI, we offer competitive compensation, generous benefits, and ample opportunities for professional growth in a dynamic environment. Join us in our mission to enhance human freedom through technology while working alongside passionate individuals who share your commitment to reliability and excellence.
StudySmarter Expert Advice🤫
We think this is how you could land Network Engineer, Reliability & Observability in London
✨Tip Number 1
Network with industry professionals! Attend meetups, webinars, or conferences related to network engineering and AI. This is a great way to make connections and learn about job openings that might not be advertised.
✨Tip Number 2
Show off your skills! Create a portfolio showcasing your projects, especially those involving reliability and observability in networks. This can really set you apart when chatting with potential employers.
✨Tip Number 3
Prepare for interviews by brushing up on your technical knowledge and incident response strategies. Be ready to discuss real-life scenarios where you've solved complex problems under pressure.
✨Tip Number 4
Don't forget to apply through our website! It’s the best way to ensure your application gets seen by the right people. Plus, we love seeing candidates who are proactive about their job search!
We think you need these skills to ace Network Engineer, Reliability & Observability in London
Some tips for your application 🫡
Tailor Your Application:Make sure to customise your CV and cover letter for the Network Engineer role. Highlight your experience in network operations and reliability engineering, and show us how your skills align with our mission at Fluidstack.
Show Your Passion:We want to see your enthusiasm for AI and networking! In your application, share why you care about improving network reliability and how you’ve tackled challenges in your previous roles. Let your passion shine through!
Be Clear and Concise:When writing your application, keep it straightforward. Use clear language and avoid jargon unless necessary. We appreciate a well-structured application that gets straight to the point while showcasing your expertise.
Apply Through Our Website:Don’t forget to submit your application through our website! It’s the best way for us to receive your details and ensures you’re considered for the role. We can’t wait to hear from you!
How to prepare for a job interview at FluidStack
✨Know Your Stuff
Make sure you brush up on your knowledge of network engineering, especially in areas like IP routing and BGP. Fluidstack is looking for someone who can demonstrate deep expertise, so be ready to discuss your hands-on experience with production networks and how you've tackled complex failures.
✨Showcase Your Problem-Solving Skills
Prepare examples of how you've led incident responses and resolved issues under pressure. Fluidstack values calmness during outages, so share specific instances where you communicated effectively with stakeholders and made critical decisions.
✨Emphasise Collaboration
Fluidstack wants a team player who can work across various teams. Be ready to talk about your experiences collaborating with deployment, operations, and software development teams. Highlight any matrix leadership roles you've held and how you built relationships to get things done.
✨Demonstrate Your Passion for Reliability
Since the role focuses on reliability and observability, express your enthusiasm for improving network quality. Discuss any relevant projects or processes you've developed that showcase your commitment to data-driven decision-making and operational excellence.