At a Glance
- Tasks: Lead reliability engineering to ensure high performance and resilience of customer-facing systems.
- Company: Join Allwyn UK, a leading lottery operator dedicated to positive societal impact.
- Benefits: Enjoy competitive salary, generous leave, health cover, and flexible benefits.
- Other info: Be part of a diverse team committed to sustainability and community support.
- Why this job: Make a real difference while working on innovative projects in a dynamic environment.
- Qualifications: Strong experience in cloud environments and hands-on engineering skills required.
The predicted salary is between 63000 - 77000 £ per year.
Description
At the heart of everything we do is our vision to change lives every day, and our mission to grow The National Lottery responsibly and champion its impact.
We are Allwyn UK, part of the Allwyn Entertainment Group – a multi-national lottery operator with a market-leading presence across the USA (Michigan and Illinois) and Europe, including Czech Republic, Austria, Greece, Cyprus and Italy.
While the main contribution of The National Lottery to society is through the funds to good causes, at Allwyn we put our purpose and values at the heart of everything we do.
Join us as we embark on a once-in-a-lifetime, largescale transformation journey by creating a National Lottery that delivers more money to good causes.
We’ll talk a bit more about us further down the page, but for now – let’s talk about the role and who we’re looking for…
A bit about the role
At Allwyn, the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate, ensuring high availability, performance, and resilience of customer-facing systems during both normal operation and peak lottery events.
The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams.
- Objectives of the role
- Own reliability outcomes across services using SLOs, SLIs, and error budgets
- Improve availability, latency, and scalability across Instant-Win and Draw-based platforms
- Lead incident response and operational readiness, including peak jackpot events
- Drive automation and platform maturity, reducing manual operational effort
- Establish clear reporting on reliability, incidents, and service health trends
- Supporting with thought-leadership and developing long-term roadmap
- What you’ll be doing
- Reliability engineering & technical leadership
- Define and govern SLOs / SLIs / error budgets across critical services
• Lead reliability design reviews across
- Web & mobile platforms
- Player Identity & Protection systems
- CMS and Geolocation services
- Drive architecture improvements for resilience (failover, degradation, scaling patterns)
- Production operations & incident leadership
- Act as incident commander for major incidents and high-severity events
• Lead 1-in-4 on-call rotation, covering
- Out-of-hours incident diagnosis
- Peak jackpot proactive monitoring
• Own end-to-end incident lifecycle
- Detection → triage → resolution → post-incident review
- Ensure blameless post-mortems with clear remediation ownership
- Observability & service insight
• Define and evolve observability strategy using
- Splunk (log analytics)
- Cloud Watch (AWS telemetry)
- Grafana (metrics visualisation)
- Quantum Metric (user behaviour insight)
• Standardise
- Alerting quality and signal-to-noise ratio
- Dashboards aligned to SLOs and customer impact
- Drive correlation between technical signals and user experience
- Automation & platform engineering
- Lead automation of operational processes using Terraform and scripting
- Improve deployment and release safety (CI/CD, progressive delivery patterns)
- Key contributor to transition strategy for ECS → EKS (Kubernetes adoption)
- Reduce operational toil through tooling, self-healing, observability and platform improvements
- Empower Level-1 operational teams with safe, controlled access to the tools they need to operate autonomously
- Capacity & performance engineering
• Own capacity planning for
- High-concurrency draw events
- Traffic spikes during jackpots
• Lead performance optimisation
- Latency reduction
- Throughput scaling
- Cost efficiency (AWS utilisation and associated log costs, observability license consumption)
- Backlog ownership & reporting
• Own and prioritise the SRE backlog, balancing
- Reliability improvements
- Technical debt
- Automation opportunities to reduce/offload toil
• Produce structured reporting covering
- SLO performance
- Incident trends and MTTR
- Platform health and risk areas
- Provide clear updates to engineering leadership and business stakeholders
- Collaboration & culture
- Embed SRE practices across engineering teams
• Mentor engineers and SREs on
- Reliability engineering
- Observability
- Incident management
- Promote a culture of automation, measurement, and continuous improvement
- Supporting with thought-leadership and developing long-term SRE roadmap for Digital Operations
- What experience we’re looking for
- Technical
- Strong experience in cloud environments, ideally AWS (ECS, with exposure or experience in EKS/Kubernetes)
• Hands-on experience with
- Terraform (Infrastructure as Code)
- Observability stacks (Splunk, Cloud Watch, Grafana)
- Strong programming skills (Python, Go, or similar)
- SRE practices
• Proven experience implementing
- SLOs, SLIs, error budgets
- Incident management frameworks
- Observability strategies
- Strong experience in distributed systems troubleshooting
- Operational
- Experience in on-call production environments
- Demonstrated leadership during high-severity incidents
Desirable Experience
- Experience migrating container platforms (ECS → EKS/Kubernetes)
- Experience supporting high-scale consumer platforms
- Familiarity with real-time analytics / customer experience tooling (e. g., Quantum Metric)
- Experience in regulated or high-availability environments
About us
At Allwyn, we are dedicated to changing lives and growing the National Lottery responsibly, championing its positive impact on people, places, and the planet.
• Innovation - We pride ourselves on it!
We’re constantly looking for new ways to excite our customers, bringing new products to market to enjoy which is all supported by our responsible play values and making them accessible to all.
- Giving back – Did you know that playing the lottery generates around £30m a week for charities and good causes in the UK?
Our aim is to have doubled this number by the end of the first 10-year license.
- Sustainability – Our aim is to become a net zero national lottery.
We have 2030 targets to decarbonise our operations and energy.
We’ve already transitioned to renewable energy providers, made our London and Watford offices zero gas, and ensured our fleet consists of low-emission vehicles.
In addition, we’re working with our value chain partners to develop a net zero target date.
- Empowering every voice – We believe in creating a culture where everyone feels they belong, can be themselves, has access to opportunities and can thrive for the benefit of good causes.
Our diverse teams are working hard to make all parts of The National Lottery inclusive – whether people play a game in a store or online, because when everyone can play, everyone wins..
An inclusive reward offering with wellbeing at the centre
At Allwyn, inclusion is built into how we care for our people.
Our benefits and policies support colleagues and their families at every stage of life and career.
By prioritising wellbeing and belonging, we create a workplace where everyone feels valued, rewarded, and empowered to succeed.
Our people are more than colleagues - they’re winners, driving positive change and making a real difference in communities.
Benefits
- Company Bonus Scheme
- Matched pension contributions up to 8.5%
- 26 days annual leave + 2 Life Days (and bank holidays)
- Single Private Health Cover
- Complimentary Private Medical
- Income Protection
- Flexible Benefits – EV Scheme, Money Coach, Will Writing, Mortgage Advice, Dental and Eye Care Schemes.
- Enhanced Family Leave (Maternity, Paternity, Adoption)
- Wellness Allowance £500
- Employee Assistance Programme
- Discounted Health Assessments
- Volunteering Days
- Matched Funding
We are a Disability Confident Leader which means we’ve taken proactive steps to ensure our workplace is accessible and inclusive for disabled and neurodivergent colleagues and candidates.
As part of this we offer an interview to disabled applicants who meet the essential requirements of the job.
If you need any assistance or adjustments to this job description or in the application process, please contact a member of the talent team at careers@allwyn. co. uk and we’ll be happy to help.
Senior / Lead Site Reliability Engineer in Watford employer: Allwyn UK
Allwyn UK in Watford is an excellent employer, offering a dynamic work culture that prioritises innovation and collaboration within the national lottery sector. Employees benefit from a competitive bonus scheme, matched pension contributions, and ample opportunities for professional growth, making it a rewarding place to advance your career in security testing.
StudySmarter Expert Advice🤫
We think this is how you could land Senior / Lead Site Reliability Engineer in Watford
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Allwyn UK or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Allwyn UK.
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Allwyn UK.
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Allwyn UK that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Senior / Lead Site Reliability Engineer in Watford
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Allwyn UK.
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Allwyn UK and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at Allwyn UK
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Allwyn UK uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.