Principal/Senior Site Reliability Engineer in City of Westminster

Principal/Senior Site Reliability Engineer in City of Westminster

City of Westminster Full-Time 60000 - 80000 £ / year (est.) Home office (partial)
R

At a Glance

  • Tasks: Design resilient cloud systems for MLOps and HPC workloads that impact science and patients.
  • Company: Join Roche, a global leader in healthcare innovation and diversity.
  • Benefits: Flexible working, relocation support, and opportunities for personal and professional growth.
  • Other info: Collaborative environment with a focus on creativity and innovation.
  • Why this job: Make a real difference in healthcare while working with cutting-edge technology.
  • Qualifications: Expertise in Infrastructure as Code and cloud-native architectures required.

The predicted salary is between 60000 - 80000 £ per year.

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

Join the Computational Sciences Center of Excellence as a Senior Site Reliability Engineer, where the platforms you build accelerate the discovery of transformative medicines. You will work alongside talented engineers in the Data & Digital Catalyst organisation to design resilient, cloud-based systems for MLOps and HPC workloads at global scale. This is a role for someone who wants their engineering craft to have real impact on science and patients.

The Opportunity:

  • You architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.
  • You design for resilience building disaster recovery and failover plans with auto-scaling and load balancing to keep critical systems available worldwide.
  • You strengthen reliability through chaos engineering running experiments that validate systems and surface weaknesses before they become incidents.
  • You build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.
  • You provide technical leadership to a team of engineers, fostering collaboration, innovation, and continuous improvement.
  • You partner across teams to align infrastructure with ML and HPC needs and to advance operational maturity through SLAs, SLOs, SLIs, and error budgets.

Who you are:

  • You bring deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or CloudFormation in AWS, Azure, or GCP for MLOps and HPC workloads.
  • You understand cloud-native and on-prem architectures including autoscaling, serverless, and multi-region deployments, and you are hands‑on with Docker, Kubernetes, and Kubeflow.
  • You are an expert in automation scripting confidently in Python, Bash, or Go, with a strong grasp of GPU‑accelerated computing and HPC workload scaling.
  • You lead through influence communicating and mentoring with clarity, and solving complex problems with a methodical approach.
  • You hold a degree in Computer Science or a related technical field or bring equivalent experience in software and site reliability engineering.

Preferred:

  • Experience with distributed ML frameworks such as Horovod or TensorFlow Distributed.
  • Familiarity with data engineering pipelines such as Apache Airflow or Apache Spark.
  • Knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.

Relocation benefits are available for this position.

At Roche Products we believe diversity drives innovation and we are committed to building a diverse and flexible working environment. All qualified applicants will receive consideration for employment without regard to race, religion or belief, sex, gender reassignment, sexual orientation, marriage and civil partnership, pregnancy and maternity, disability or age. We recognise the importance of flexible working and will review all applicants’ requests with care. At Roche difference is valued and we are proud to be an equal opportunity employer where you are encouraged to bring your whole self to work.

Principal/Senior Site Reliability Engineer in City of Westminster employer: Roche Holding AG

At Roche, we pride ourselves on fostering a culture that values individuality and encourages open dialogue, making it an exceptional place to work. As a Disease Level Partner in South Central UK, you will have the opportunity to make a meaningful impact on patient outcomes while enjoying a supportive environment that promotes personal and professional growth. With a commitment to diversity and flexible working arrangements, Roche ensures that every employee feels valued and empowered to thrive.

R

Contact Details:

Roche Holding AG Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Principal/Senior Site Reliability Engineer in City of Westminster

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Roche Holding AG or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Roche Holding AG.

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Roche Holding AG.

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Roche Holding AG that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Principal/Senior Site Reliability Engineer in City of Westminster

Infrastructure as Code
Terraform
Pulumi
CloudFormation
AWS
Azure
GCP

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Roche Holding AG.

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Roche Holding AG and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at Roche Holding AG

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Roche Holding AG uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.