At a Glance
- Tasks: Identify reliability risks and drive initiatives to enhance system performance.
- Company: Join a forward-thinking tech company focused on innovation and collaboration.
- Benefits: Enjoy competitive pay, health perks, remote work options, and growth opportunities.
- Other info: Dynamic team environment with opportunities for professional development and career advancement.
- Why this job: Make a real impact by improving system reliability and working with cutting-edge technologies.
- Qualifications: 4+ years in software engineering with SRE or operations experience required.
The predicted salary is between 63000 - 77000 Β£ per year.
Responsibilities:
- Proactively identify reliability risks and independently drive initiatives to address them before they become incidents.
- Participate in and continuously improve our on-call rotation, including incident response, triage, and leading blameless post-incident reviews.
- Define and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across services.
- Scope technical projects and break them down into user stories and tasks, driving them to completion with minimal oversight.
- Make sound technical decisions, leveraging input from teammates and contributing to technical conversations across engineering teams.
- Automate the provisioning and management of infrastructure using Infrastructure as Code (IaC) tools such as Terraform.
You may be a good fit if:
- You have at least 4 years of experience working in a professional environment as a Software Engineer (with some SRE or operations responsibilities).
- You have contributed to the design and build of cloud-native applications written in Python, Java or Go.
- You have strong hands-on experience with observability.
- You understand the difference between monitoring and observability, and can articulate how metrics, logs, and traces work together.
- You have worked with Infrastructure As Code tooling, for example Terraform.
- You have participated in on-call rotations and are comfortable leading incident response under pressure, communicating clearly with stakeholders throughout.
- You are comfortable taking ownership of initiatives or projects independently, from scoping through to delivery.
- You build effective working relationships, give and receive constructive feedback openly, and are trusted by colleagues at all levels.
Technologies we use include:
- Python, Java, and Go are our primary server languages.
- Our browser applications are based on Angular and React.
- Code lives in GitHub and flows to production through a CI/CD pipeline built on GitHub Actions, with some workloads on Jenkins.
- Infrastructure runs on AWS (EC2) with workloads on Kubernetes-managed Docker containers.
- Datadog is our primary observability platform β experience with Datadog APM, dashboards, monitors, and RUM is a plus.
- Infrastructure is managed as code using Terraform.
Site Reliability Engineer in London employer: Infinity Quest
As a Senior SAP Consultant at our company, you will thrive in a dynamic and supportive remote work environment that prioritises employee growth and development. We offer competitive benefits, a collaborative culture, and opportunities to lead innovative projects that make a real impact on our clients' success. Join us to advance your career while working with cutting-edge technologies and a team of dedicated professionals committed to excellence.
We think you need these skills to ace Site Reliability Engineer in London
Reliability Risk Identification
Incident Response
Monitoring Strategies
Logging Strategies
Distributed Tracing
SLO/SLI/SLA Management
Technical Project Scoping