Site Reliability Engineer I in London

Site Reliability Engineer I in London

London Full-Time 63000 - 77000 Β£ / year (est.) Home office (partial)
Slide to start your application
Start application
T

At a Glance

  • Tasks: Join us to enhance platform reliability and support incident response in a dynamic environment.
  • Company: Trainline, Europe's leading independent rail platform focused on greener travel.
  • Benefits: Enjoy private healthcare, generous leave, and a work-from-abroad policy.
  • Other info: Hybrid working model with clear career paths and personal learning budgets.
  • Why this job: Make a real impact on sustainable travel while developing your tech skills.
  • Qualifications: Experience with AWS, Linux, and scripting; a growth mindset is essential.

The predicted salary is between 63000 - 77000 Β£ per year.

About us At Trainline, our purpose is to empower greener travel choices, connecting people and places. Trainline enables millions of travellers to find and book the best value tickets across carriers, fares, and journey options through our highly rated mobile app, website, and B2B partner channels. Great journeys start with Trainline. We're Europe's leading independent rail platform, helping millions of travellers find and book the best-value rail and coach journeys across our app, website and partner channels. Our job is to make the green travel choice the best choice. By building a better train travel experience, we help more people choose rail - creating a positive impact for customers, our business and the planet. Now is a brilliant time to join us and help shape the future of travel.

Introducing Reliability & Operations Engineering Trainline is a fast-growing tech company powering world-class digital journeys for millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices. The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery, respond to incidents, and continuously strengthen system reliability.

We're looking for a mid-level Site Reliability Engineer to help drive this forward. You'll bring solid production experience, a growth mindset, and a willingness to challenge and be challenged β€” contributing to platform reliability while developing broader technical ownership with support from senior engineers.

  • Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform
  • Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration
  • Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience
  • Taking part in the SRE on-call rotation
  • Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis
  • Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD)
  • Ensuring relevant operational data is surfaced quickly and clearly during live incidents
  • Making informed tooling and technology choices using SRE principles, balancing team and business needs
  • Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling
  • Collaborating with product engineering teams to ensure services are operationally ready and deployed safely
  • Advising on reliability and resilience practices
  • Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals
  • Prioritising work effectively and collaborating using agile processes to deliver against team and business goals

Our Tech Stack: AWS, New Relic, ELK stack, Grafana

  • Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar
  • Experience working with cloud providers (preferably AWS)
  • Experience troubleshooting Linux operating systems
  • Experience of scripting in at least one language (preferably Python)
  • Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts
  • Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions
  • Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform

Enjoy fantastic perks like private healthcare & dental insurance, a generous work from abroad policy, 2-for-1 share purchase plans, an EV Scheme to further reduce carbon emissions, extra festive time off, and excellent family-friendly benefits. We prioritise career growth with clear career paths, transparent pay bands, personal learning budgets, and regular learning days. We're operating a hybrid model and ask that Trainliners work from the office a minimum of 60% of their time over a 12-week period. We also have a 28-day Work from Abroad policy.

Think Big - We're building the future of rail
Own It - We focus on every customer, partner and journey
Travel Together - We're one team
Do Good - We make a positive impact

We know that having a diverse team makes us better and helps us succeed. And we mean all forms of diversity - gender, ethnicity, sexuality, disability, nationality and diversity of thought.

Site Reliability Engineer I in London employer: Trainline

Trainline is an exceptional employer, dedicated to fostering a culture of innovation and sustainability in the travel industry. With a strong emphasis on employee growth, we offer clear career paths, personal learning budgets, and a supportive environment for mentorship and collaboration. Our hybrid work model, generous benefits, and commitment to diversity make Trainline a rewarding place to build a meaningful career while contributing to a greener future.

T

Contact Details:

Trainline Recruitment Team

We think you need these skills to ace Site Reliability Engineer I in London

AWS
Cloud-Native Architecture
CI/CD Pipelines
DevOps Practices
Incident Management
Database Reliability
Observability Tooling