Senior Site Reliability Engineer

Senior Site Reliability Engineer

Full-Time 60000 - 80000 £ / year (est.) Working from home possible
Runware

At a Glance

  • Tasks: Ensure reliability and performance of critical production services while improving observability and reducing incidents.
  • Company: Join Runware, a remote-first tech collective powering the world's intelligence.
  • Benefits: Enjoy flexible hours, generous paid time off, and meaningful stock options.
  • Other info: Collaborative environment with opportunities for personal growth and exciting company retreats.
  • Why this job: Make a real impact on cutting-edge AI infrastructure and work with innovative technologies.
  • Qualifications: Strong experience in SRE or similar roles, with skills in distributed systems and automation.

The predicted salary is between 60000 - 80000 £ per year.

Runware is building high-performance infrastructure and products to power the world's intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.

What you’ll do:

  • Own and improve the reliability, availability and performance of critical production services across the Runware platform.
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation.
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements.
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience.
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows.

Requirements:

  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role.
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure.
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing.
  • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil.
  • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP.
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation.

Bonus:

  • Experience operating high-throughput or low-latency APIs and distributed systems.
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads.
  • Experience with RabbitMQ or other distributed messaging and queueing systems.
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems.
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments.
  • Experience building automated scaling, capacity management or self-healing systems.

Benefits:

  • We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time.
  • We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.
  • Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.
  • Generous paid time off – vacation, sick days, public holidays.
  • Meaningful stock options – share in the upside you create.
  • Remote-first setup – work from home anywhere we can employ you.
  • Flexible hours – own your schedule outside core collaboration blocks.
  • Family leave – paid maternity, paternity, and caregiver time.
  • Company retreats – twice-yearly gatherings in inspiring locations.

Senior Site Reliability Engineer employer: Runware

Runware is an exceptional employer for those looking to lead in the DevOps space, offering a remote-first work culture that prioritises flexibility and work-life balance. With generous paid time off, meaningful stock options, and opportunities for professional growth through mentorship, employees are empowered to thrive in a fast-paced environment while enjoying the benefits of collaborative retreats twice a year. Join us to shape innovative infrastructure solutions that drive AI delivery globally, all while working from anywhere in the UK.

Runware

Contact Details:

Runware Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Senior Site Reliability Engineer

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Runware or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Runware.

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Runware.

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Runware that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Senior Site Reliability Engineer

Site Reliability Engineering
Production Systems Troubleshooting
Distributed Systems Debugging
Observability Systems Design
SLIs and SLOs Understanding
Incident Management
Kubernetes

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Runware.

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Runware and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at Runware

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Runware uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.