Site Reliability Engineer

Site Reliability Engineer

Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
intro

At a Glance

  • Tasks: Take ownership of production incidents and troubleshoot complex issues across various systems.
  • Company: Fast-growing FinTech revolutionising banking technology for financial institutions.
  • Benefits: Competitive salary, discretionary bonus, and compensated on-call support.
  • Other info: Dynamic startup environment with opportunities for growth and innovation.
  • Why this job: Make a real impact by improving production processes and working with cutting-edge tech.
  • Qualifications: Experience in Production Support and knowledge of AWS or GCP required.

The predicted salary is between 63000 - 77000 Β£ per year.

  • Site Reliability Engineer/Production Support Engineer | Fin Tech |
  • London | 4 Days Onsite

Are you the engineer people turn to when production goes wrong?

We're partnering with a fast-growing Fin Tech building mission-critical banking technology for regulated financial institutions.

As the platform continues to scale, they're looking for a

Site reliability engineer / Production Support Engineer who can take real ownership of production while working across the boundaries of

Support, SRE, Dev Ops and Software Engineering .

This is a hands-on production role at its core.

You'll be responsible for investigating incidents, restoring services and identifying root causes, but you'll also have the opportunity to go deeper, whether that's improving observability, working with cloud infrastructure, automating processes or getting into the backend code to implement smaller fixes and improvements.

Being a growing Fin Tech means responsibilities aren't confined to one box.

They're looking for someone who enjoys wearing multiple hats and wants genuine ownership over how production is supported and improved.

  • What you'll be doing
  • Taking ownership of production incidents end-to-end , from initial investigation and incident response through to resolution and post-incident review.
  • Troubleshooting complex issues across applications, infrastructure, databases, APIs and distributed systems .
  • Performing

Root Cause Analysis (RCA) and helping implement permanent fixes and preventative improvements.

  • Working within structured production processes including

ITIL, change management, CAB and incident/problem management .

  • Improving monitoring, alerting and observability to identify issues before they impact customers.
  • Working hands-on with

AWS or GCP environments and collaborating closely with SRE/Dev Ops and Platform teams.

  • Reading and debugging backend code , making smaller bug fixes, modifications and production improvements where appropriate.
  • Automating repetitive operational processes and improving the efficiency of production support.
  • Supporting releases and deployments into production.
  • Working closely with Software Engineering to resolve larger or more complex application issues.
  • Participating in a compensated out-of-hours on-call rota .

What we're looking for

  • Strong experience within

Production Support, Application Support or Production Engineering .

  • Experience operating at

L2/L3 level , with the ability to investigate complex technical issues rather than simply escalating them.

  • Strong knowledge of incident management, incident response, RCA, change management and production support processes .
  • Hands-on experience with

AWS or GCP .

  • Exposure to

SRE/Dev Ops practices , CI/CD and modern production environments.

  • Experience with monitoring and observability tooling such as

Grafana, Prometheus, Datadog, Splunk, New Relic or similar .

  • Comfortable reading, debugging and modifying backend code . The specific programming language isn't important, Python, Java, Go, C#, Kotlin or similar are all relevant.
  • Experience supporting distributed, event-driven or highly available systems .
  • Comfortable working closely with Software Engineering, Dev Ops/SRE and Infrastructure teams.
  • Nice to have
  • Fin Tech, banking, payments, trading or wider financial services experience.
  • Terraform or other

Infrastructure as Code tooling.

  • Kubernetes and Docker.
  • Experience within a startup or scale-up environment .
  • Experience supporting high-volume or business-critical transactional platforms.
  • London – 4 days onsite
  • Β£55,000–£85,000 + discretionary bonus + benefits + compensated on-call

If you're a Production Support Engineer who enjoys going beyond the ticket queue, getting into the infrastructure, understanding the code and improving how production operates, I'd be keen to speak.

#J-18808-Ljbffr

Site Reliability Engineer employer: intro

As a leading FinTech company based in Central London, we pride ourselves on fostering a dynamic and innovative work culture that empowers our employees to thrive. With a strong focus on professional development, we offer numerous growth opportunities and encourage collaboration across teams, ensuring that every voice is heard. Our commitment to leveraging cutting-edge technology, such as AI tools, not only enhances productivity but also makes working here both meaningful and rewarding.

intro

Contact Details:

intro Recruitment Team

We think you need these skills to ace Site Reliability Engineer

Production Support
Incident Management
Incident Response
Root Cause Analysis (RCA)
Change Management
AWS
GCP