Site Reliability Engineer in London

Site Reliability Engineer in London

London Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
I

At a Glance

  • Tasks: Identify reliability risks and drive initiatives to enhance system performance.
  • Company: Join a forward-thinking tech company focused on innovation and collaboration.
  • Benefits: Enjoy competitive pay, health perks, remote work options, and growth opportunities.
  • Other info: Dynamic team environment with opportunities for professional development and career advancement.
  • Why this job: Make a real impact by improving system reliability and working with cutting-edge technologies.
  • Qualifications: 4+ years in software engineering with SRE or operations experience required.

The predicted salary is between 63000 - 77000 Β£ per year.

Responsibilities:

  • Proactively identify reliability risks and independently drive initiatives to address them before they become incidents.
  • Participate in and continuously improve our on-call rotation, including incident response, triage, and leading blameless post-incident reviews.
  • Define and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across services.
  • Scope technical projects and break them down into user stories and tasks, driving them to completion with minimal oversight.
  • Make sound technical decisions, leveraging input from teammates and contributing to technical conversations across engineering teams.
  • Automate the provisioning and management of infrastructure using Infrastructure as Code (IaC) tools such as Terraform.

You may be a good fit if:

  • You have at least 4 years of experience working in a professional environment as a Software Engineer (with some SRE or operations responsibilities).
  • You have contributed to the design and build of cloud-native applications written in Python, Java or Go.
  • You have strong hands-on experience with observability.
  • You understand the difference between monitoring and observability, and can articulate how metrics, logs, and traces work together.
  • You have worked with Infrastructure As Code tooling, for example Terraform.
  • You have participated in on-call rotations and are comfortable leading incident response under pressure, communicating clearly with stakeholders throughout.
  • You are comfortable taking ownership of initiatives or projects independently, from scoping through to delivery.
  • You build effective working relationships, give and receive constructive feedback openly, and are trusted by colleagues at all levels.

Technologies we use include:

  • Python, Java, and Go are our primary server languages.
  • Our browser applications are based on Angular and React.
  • Code lives in GitHub and flows to production through a CI/CD pipeline built on GitHub Actions, with some workloads on Jenkins.
  • Infrastructure runs on AWS (EC2) with workloads on Kubernetes-managed Docker containers.
  • Datadog is our primary observability platform β€” experience with Datadog APM, dashboards, monitors, and RUM is a plus.
  • Infrastructure is managed as code using Terraform.

Site Reliability Engineer in London employer: Infinity Quest

As a Senior SAP Consultant at our company, you will thrive in a dynamic and supportive remote work environment that prioritises employee growth and development. We offer competitive benefits, a collaborative culture, and opportunities to lead innovative projects that make a real impact on our clients' success. Join us to advance your career while working with cutting-edge technologies and a team of dedicated professionals committed to excellence.

I

Contact Details:

Infinity Quest Recruitment Team

We think you need these skills to ace Site Reliability Engineer in London

Reliability Risk Identification
Incident Response
Monitoring Strategies
Logging Strategies
Distributed Tracing
SLO/SLI/SLA Management
Technical Project Scoping