Site Reliability Engineer in London

Site Reliability Engineer in London

London Full-Time 63000 - 77000 £ / year (est.) No working from home possible
Willis Re (UK) Limited

At a Glance

  • Tasks: Design, operate, and automate Azure platforms for high availability and performance.
  • Company: Join Willis Re, a forward-thinking company in the reinsurance industry.
  • Benefits: Flexible work environment, diverse culture, and opportunities for professional growth.
  • Other info: Be part of a dynamic team focused on cutting-edge analytical solutions.
  • Why this job: Make a real impact by enhancing service reliability and driving innovation.
  • Qualifications: Experience with Azure, Terraform, and DevOps practices is essential.

The predicted salary is between 63000 - 77000 £ per year.

Willis Re is expanding its Global business in 2026 and the Cloud and Infrastructure team will need to support this growth by designing, provisioning then supporting the platforms to enable this. We are seeking an experienced Site Reliability Engineer (SRE) to join our Cloud & Infrastructure team. The successful candidate will be responsible for designing, operating, automating, and continuously improving enterprise-scale Azure platforms, ensuring high availability, resiliency, security, and performance. The role combines software engineering, cloud architecture, infrastructure automation, and operational excellence to improve service reliability and reduce operational overhead through automation and engineering best practices. The ideal candidate will have strong experience with Microsoft Azure, DevOps, Terraform, API, disaster recovery planning, and enterprise-scale resilience engineering.

Key Responsibilities

  • Platform Reliability & Operations
    • Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
    • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
    • Proactively monitor platform health and performance using observability tooling.
    • Perform root cause analysis and implement permanent fixes for recurring incidents.
    • Participate in incident management and on-call support rotations where required.
    • Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
    • Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
    • Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.
  • Azure Cloud Engineering
    • Design, deploy, and manage Azure infrastructure services including: Virtual Networks, Application Gateways, API Management, Azure Kubernetes Service (AKS), Azure Firewall, Azure Storage, Key Vault, Azure Monitor, Azure AI Services.
    • Implement cloud platform standards and best practices.
    • Support multi-region Azure deployments and platform modernisation initiatives.
    • Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.
  • Infrastructure as Code (Terraform)
    • Develop and maintain Terraform modules and reusable infrastructure patterns.
    • Implement Infrastructure as Code (IaC) standards and governance controls.
    • Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
    • Manage Terraform state securely and consistently across environments.
  • Azure DevOps & Automation
    • Build and maintain Azure DevOps CI/CD pipelines.
    • Automate infrastructure provisioning and application deployments.
    • Implement testing, security scanning, policy compliance, and release gates.
    • Support GitOps and platform engineering practices.
    • Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
  • Resilience, Disaster Recovery & Failover
    • Design and implement highly available Azure architectures.
    • Develop and maintain disaster recovery and business continuity capabilities.
    • Implement and test regional failover strategies, active/passive architectures, active/active deployments, Traffic Manager and Front Door failover patterns, database resiliency and replication, backup and recovery solutions.
    • Conduct regular resilience and recovery testing exercises.
    • Identify and reduce single points of failure across platforms.
    • Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.
  • Security & Governance
    • Ensure platforms are secure-by-design.
    • Work closely with Security and Architecture teams to implement Zero Trust principles, RBAC controls / Managed Identities, network segmentation, secrets management.
    • Support compliance requirements and operational audits.
    • Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors.
    • Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.
  • Continuous Improvement
    • Drive automation and reduction of manual operational tasks.
    • Improve deployment reliability and platform observability.
    • Contribute to architecture standards, runbooks, and operational documentation.
    • Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.

About You

At Willis Re we work in a fast-paced evolving environment with a growth mindset and outcome focus.

Core Skills & Experience

  • Proven experience in senior positions.
  • Ability to articulate complex ideas and scenarios into business language.
  • Ability to manage and prioritise workload.
  • Worked in a fast-paced complex environment and demonstrate hands-on experience.
  • Microsoft Azure Expert.
  • Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
  • Hands-on proficiency with infrastructure-as-code and automated deployments (Must have Terraform and Azure DevOps).
  • Strong understanding of Azure security controls, governance, and compliance frameworks.
  • Strong FinOps expertise.
  • Scripting skills (PowerShell, Bash, or Python).
  • Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
  • Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.

Nice to have

  • Experience with Observability tools such as DataDog.
  • Understanding of Zero Trust Architecture.
  • Snowflake experience.
  • GitHub Enterprise.

About Willis Re

We combine specialist broking with analytics, modeling and research to help insurers optimize risk transfer, strengthen balance sheets and achieve sustainable growth. Our approach is relationship-driven, transparent and outcome-focused. At the heart of Willis Re is a focus on delivering the most cutting-edge analytical solutions to enable more informed, better decision-making for risk selection, portfolio optimization, and capital management. Willis Re is committed to embracing a diverse, inclusive, and flexible work environment. We provide equal opportunity to all qualified individuals regardless of race, colour, religion, age, gender, gender expression, national origin, veteran status, disability, orientation and any other legally protected categories.

Site Reliability Engineer in London employer: Willis Re (UK) Limited

Willis Re is an exceptional employer that fosters a collaborative and innovative work culture, where Capital Advisory Specialists can thrive. With a strong emphasis on professional development, employees are encouraged to enhance their skills through diverse projects and cross-functional teamwork, all while enjoying the dynamic environment of the insurance capital markets. The company's commitment to integrity and excellence ensures that team members are well-supported in delivering high-quality advisory solutions to clients.

Willis Re (UK) Limited

Contact Details:

Willis Re (UK) Limited Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Site Reliability Engineer in London

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Willis Re (UK) Limited or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Willis Re (UK) Limited.

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Willis Re (UK) Limited.

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Willis Re (UK) Limited that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Site Reliability Engineer in London

Microsoft Azure
DevOps
Terraform
API Management
Disaster Recovery Planning
Infrastructure Automation
Site Reliability Engineering Principles

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Willis Re (UK) Limited.

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Willis Re (UK) Limited and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at Willis Re (UK) Limited

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Willis Re (UK) Limited uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.