At a Glance
- Tasks: Oversee production support, manage incidents, and build automation for reliability.
- Company: Join Monument, the UK's fastest growing fintech, transforming banking for the mass affluent.
- Benefits: Hybrid work, competitive salary, and the chance to make a real impact.
- Other info: Dynamic team culture focused on collaboration and continuous improvement.
- Why this job: Own production reliability at a pre-IPO bank and work with cutting-edge AI tools.
- Qualifications: Experience in SRE or production support, with strong incident response skills.
The predicted salary is between 63000 - 77000 £ per year.
hackajob is partnering directly with Monument to hire for this role. Location: London (Oxford Circus) | Hybrid: 2 days per week | Reports to Head of Cloud Operations
ABOUT MONUMENT
We're building something genuinely rare: a financial brand designed for the mass affluent, the professionals, entrepreneurs and ambitious savers that traditional banks have systematically underserved for decades. We exist to make managing wealth simpler, smarter and more human, treating every client's wealth with the same care as if it were our own. We hold over £7 billion in client savings, serve more than 100,000 clients, and were named the UK's fastest growing fintech in 2025. The momentum is real.
THE OPPORTUNITY
Monument's production environment is the heartbeat of a licensed bank, and the SRE role is the single point of ownership when incidents occur. You will directly oversee the offshore Production Support team, run on-call and incident response, and ensure fast detection, triage and restoration of services. This is not a passive monitoring role. You are expected to understand at a working level all of Monument's key system flows from the services, partners and teams in play and to actively debug incidents, escalate effectively and drive permanent fixes. You will also be a builder: using AI tools for automated alert correlation, root cause analysis and runbook generation. For the right person, this is a rare opportunity to own production reliability at a pre-IPO challenger bank, operating at the intersection of deep engineering and real commercial consequence.
WHAT YOU'LL DO
- Directly oversee the offshore Production Support team and be the single point person when incidents occur, escalating only to Head Of when required.
- Run on-call and incident response; ensure fast detection, triage, and restoration.
- Maintain observability standards (logs, metrics, traces) and alert quality (low noise, high signal).
- Understand at a working level all key system flows, the services, partners, and teams in play, and how to actively debug an incident.
- Lead reliability engineering: resilience patterns, performance tuning, capacity planning.
- Facilitate post-incident reviews and track actions to completion.
- Use AI tools for automated alert correlation, root cause analysis, and runbook generation.
- Hunt for routine/common tasks and formulate plans on how to automate and then execute them.
THE MINDSET
- You’ll thrive here if you live by the same principles that define all Monument builders:
- Ownership under pressure - when things go wrong, you are calm, decisive, and effective. You own the incident until it's resolved.
- Builder - you don't just respond to incidents; you build the automation that prevents them or resolves them faster next time.
- Deeply curious - you understand the full system landscape and how services interact. You can debug across layers.
- Automation-first - every manual task is a candidate for automation. You actively hunt for toil and eliminate it.
- Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source.
WHAT YOU BRING
- Strong SRE or production support experience with accountability for incident response in a production environment.
- Deep understanding of observability tools, alerting, logging, and distributed systems debugging.
- Experience managing and working with offshore support teams.
- Hands-on experience with reliability engineering: resilience patterns, performance tuning, capacity planning.
- Active use of AI tools for incident triage, automation, and runbook generation.
- Ability to understand complex system flows across multiple services and third-party integrations.
- Experience in financial services or similarly regulated environments is a strong advantage.
WHAT'S IN IT FOR YOU
Be the person who keeps Monument running - your work directly protects clients and the business. Build automation that genuinely matters: every runbook you automate and every alert you tune makes the system more resilient. Work with modern AI tools as a core part of your daily workflow, not as a novelty. Own production reliability at a critical stage of Monument's growth, with real responsibility and real impact.
OUR VALUES
At Monument, our values shape how we make decisions, how we treat each other when things get hard, and how we show up for clients who expect more than standard banking. We set ambitious goals and hold ourselves to them, not because it looks good, but because our clients' outcomes depend on it. When something isn't working, we say so early, learn from it, and move. We don't wait for perfect conditions, and we don't protect egos over progress. We work as a genuine team, which means real collaboration, honest conversations when we disagree, and shared accountability when things go wrong. We know better decisions come from different perspectives, so we actively value the range of experiences and backgrounds our people bring. We're always asking whether there's a smarter way to do what we do, not for the sake of change, but because standing still isn't an option in the market we're in. If that sounds like how you like to work, we'd like to hear from you.
Site Reliability Engineer / Production Support in London employer: Hackajob Ltd
At loveholidays, we pride ourselves on fostering a collaborative and innovative work culture that empowers our employees to thrive. As a Product Designer, you'll have the opportunity to contribute to meaningful projects that enhance customer experiences while enjoying a range of benefits, including professional development opportunities and a supportive team environment in a vibrant location. Join us in our mission to make travel accessible for everyone and be part of a company that values your creativity and input.
StudySmarter Expert Advice🤫
We think this is how you could land Site Reliability Engineer / Production Support in London
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Hackajob Ltd or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Hackajob Ltd.
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Hackajob Ltd.
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Hackajob Ltd that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Site Reliability Engineer / Production Support in London
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Hackajob Ltd.
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Hackajob Ltd and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at Hackajob Ltd
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Hackajob Ltd uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.