All potential applicants are encouraged to scroll through and read the complete job description before applying.
hackajob is partnering directly with Monument to hire for this role.
Site Reliability Engineer/ Production Support
Location London (Oxford Circus) | Hybrid: 2 days per week | Reports to Head of Cloud Operations
ABOUT MONUMENT
We're building something genuinely rare: a financial brand designed for the mass affluent, the professionals, entrepreneurs and ambitious savers that traditional banks have systematically underserved for decades.
We exist to make managing wealth simpler, smarter and more human, treating every client's wealth with the same care as if it were our own.
We hold over Β£7 billion in client savings, serve more than 100,000 clients, and were named the UK's fastest growing fintech in 2025. The momentum is real.
THE OPPORTUNITY
Monuments production environment is the heartbeat of a licensed bank, and the SRE role is the single point of ownership when incidents occur. You will directly oversee the offshore Production Support team, run on-call and incident response, and ensure fast detection, triage and restoration of services.
This is not a passive monitoring role. You are expected to understand at a working level all of Monuments key system flows from the services, partners and teams in play and to actively debug incidents, escalate effectively and drive permanent fixes. You will also be a builder: using AI tools for automated alert correlation, root cause analysis and runbook generation.
For the right person, this is a rare opportunity to own production reliability at a pre-IPO challenger bank, operating at the intersection of deep engineering and real commercial consequence.
WHAT YOU'LL DO
- Directly oversee the offshore Production Support team and be the single point person when incidents occur, escalating only to Head Of when required.
- Run on-call and incident response; ensure fast detection, triage, and restoration.
- Maintain observability standards (logs, metrics, traces) and alert quality (low noise, high signal).
- Understand at a working level all key system flows, the services, partners, and teams in play, and how to actively debug an incident. xgikmsk
- Lead reliability engineering: resilience patterns, performance tuning, capacity planning.
- Facilitate post-incident Please click on the apply button to read the full job description
Site Reliability Engineer / Production Support in Liverpool employer: Hackajob
Joining Google as a Security Platform Engineer in the UK Public Sector means becoming part of a dynamic and innovative team dedicated to delivering secure private cloud services for critical customers. With a strong emphasis on employee growth, you will have access to cutting-edge technology and collaborative opportunities that foster professional development. The inclusive work culture at Google encourages creativity and teamwork, making it an exceptional employer for those seeking meaningful and impactful work in a supportive environment.