At a Glance
- Tasks: Lead the design and performance of our AI inference platform, ensuring low latency and high reliability.
- Company: Join a remote-first tech collective focused on innovation and collaboration.
- Benefits: Enjoy flexible hours, generous paid time off, and meaningful stock options.
- Other info: Opportunity for mentorship and growth in a fast-paced environment.
- Why this job: Make a real impact on cutting-edge AI technology while working remotely.
- Qualifications: Strong backend development skills and experience with high-performance systems.
The predicted salary is between 48000 - 84000 £ per year.
We are looking for a Staff Engineer to take technical ownership of latency, throughput, and reliability across Runware's AI inference platform. This is a senior technical leadership role for someone who obsesses over performance at scale, from request ingress through GPU execution to result delivery, and who can consistently turn ambitious targets such as sub-one-second inference into production reality. As a Staff Engineer, you will define and drive the architecture, standards, and execution needed to make Runware one of the fastest and most reliable inference platforms in the market. You will work deeply across backend services, distributed systems, GPU workloads, and infrastructure, partnering closely with product, ML, and platform teams. This role is ideal for someone who enjoys operating at the intersection of systems design, performance engineering, and real-world scale, and who wants clear ownership over outcomes that matter directly to customers.
What You'll Do
- Own end-to-end inference performance across the platform, with clear responsibility for latency, throughput, and reliability targets.
- Lead the architecture and design of core inference systems, including request routing, async execution, queuing, GPU scheduling, and result delivery.
- Drive the platform toward sub-1 second inference where feasible, identifying bottlenecks across networking, services, storage, and GPU execution.
- Make high-impact architectural decisions with performance, scalability, and operational simplicity as first-class concerns.
- Partner with ML and model teams to ensure models are production-ready from a performance perspective (cold starts, batching, memory usage, concurrency).
- Define performance budgets, SLAs, and success metrics, and ensure they are measured, visible, and actively improved.
- Lead deep-dive investigations into latency spikes, throughput degradation, and system-level performance issues.
- Influence and mentor engineers across teams on performance engineering, distributed systems thinking, and operational excellence.
- Improve tooling, observability, and profiling capabilities to make performance issues easier to detect and reason about.
- Advocate for pragmatic engineering best practices around testing, benchmarking, rollouts, and documentation.
Requirements
- Excellent experience in software engineering, with a strong focus on backend and systems development (PHP, Python, Go, Rust, or similar).
- Proven experience building and operating high-performance, low-latency distributed systems in production.
- Deep understanding of asynchronous processing, queues, concurrency models, and back pressure.
- Strong intuition for performance trade-offs across CPU, GPU, networking, storage, and application layers.
- Experience making and defending critical architectural decisions in complex systems.
- Hands-on experience troubleshooting real production issues under load (latency, saturation, cascading failures).
- Familiarity with modern cloud infrastructure, CI/CD, and observability stacks (metrics, tracing, profiling).
- Ability to communicate clearly and influence across teams in a remote-first environment.
- Strong mentorship mindset and a desire to raise the technical bar across the organisation.
Nice to have
- Experience working on AI/ML inference platforms, GPU-backed workloads, or performance-critical compute systems.
- Knowledge of model optimisation techniques (batching, quantisation, warm-starts, memory management).
- Experience with infrastructure-as-code and DevOps practices.
- Background in startups or fast-paced environments where speed, ownership, and pragmatism matter.
- Prior ownership of latency or throughput SLOs at scale.
Benefits
- We are a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time.
- We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.
- Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.
- Generous paid time off - vacation, sick days, public holidays.
- Meaningful stock options - share in the upside you create.
- Remote-first setup - work from home anywhere we can employ you.
- Flexible hours - own your schedule outside core collaboration blocks.
- Family leave - paid maternity, paternity, and caregiver time.
- Company retreats - twice-yearly gatherings in inspiring locations.
Please note: We are unable to offer visa sponsorship in the UK at this time. Candidates must have existing right to work in the UK.
Staff Software Engineer - Inference & Performance employer: Runware
Runware is an exceptional employer for those looking to lead in the DevOps space, offering a remote-first work culture that prioritises flexibility and work-life balance. With generous paid time off, meaningful stock options, and opportunities for professional growth through mentorship, employees are empowered to thrive in a fast-paced environment while enjoying the benefits of collaborative retreats twice a year. Join us to shape innovative infrastructure solutions that drive AI delivery globally, all while working from anywhere in the UK.
StudySmarter Expert Advice🤫
We think this is how you could land Staff Software Engineer - Inference & Performance
✨Tip Number 1
Network like a pro! Reach out to folks in your industry on LinkedIn or at meetups. A personal connection can often get your foot in the door faster than a CV.
✨Tip Number 2
Prepare for those technical interviews! Brush up on your coding skills and system design principles. We recommend doing mock interviews with friends or using platforms that simulate real interview scenarios.
✨Tip Number 3
Showcase your projects! Whether it's on GitHub or a personal website, having a portfolio of your work can really impress hiring managers. It’s a great way to demonstrate your skills and passion.
✨Tip Number 4
Don’t forget to apply through our website! It’s the best way to ensure your application gets seen by the right people. Plus, we love seeing candidates who are proactive about their job search!
We think you need these skills to ace Staff Software Engineer - Inference & Performance
Some tips for your application 🫡
Show Your Passion for Performance:When writing your application, let us see your obsession with performance at scale. Share specific examples of how you've tackled latency and throughput challenges in your previous roles. We want to know how you can turn ambitious targets into reality!
Be Clear and Concise:Keep your application straightforward and to the point. Highlight your relevant experience in backend and systems development without fluff. We appreciate clarity, so make sure your skills and achievements shine through without unnecessary jargon.
Tailor Your Application:Make sure to customise your application to align with our job description. Mention your experience with distributed systems, GPU workloads, and any performance engineering you've done. This shows us that you understand what we're looking for and how you fit into our team.
Apply Through Our Website:We encourage you to apply directly through our website. It’s the best way for us to receive your application and ensures you’re considered for the role. Plus, it gives you a chance to explore more about our culture and values while you're at it!
How to prepare for a job interview at Runware
✨Know Your Tech Inside Out
Make sure you’re well-versed in the technologies mentioned in the job description, like PHP, Python, Go, or Rust. Brush up on your knowledge of distributed systems and performance engineering principles, as you'll need to demonstrate a deep understanding of these areas during the interview.
✨Prepare for Performance Scenarios
Think about real-world scenarios where you've tackled latency or throughput issues. Be ready to discuss specific examples of how you identified bottlenecks and what architectural decisions you made to improve performance. This will show your hands-on experience and problem-solving skills.
✨Showcase Your Mentorship Skills
Since this role involves influencing and mentoring other engineers, prepare to talk about your past experiences in guiding teams. Share examples of how you've raised the technical bar and fostered a culture of excellence in previous roles.
✨Ask Insightful Questions
Prepare thoughtful questions that reflect your interest in the company’s goals and challenges. Inquire about their current performance metrics, architectural decisions, or how they handle scaling issues. This not only shows your enthusiasm but also helps you gauge if the company is the right fit for you.