Site Reliability Engineer

Site Reliability Engineer

Full-Time 28800 - 48000 £ / year (est.) Home office (partial)
B

At a Glance

  • Tasks: Enhance system reliability and performance while resolving incidents with a strong engineering approach.
  • Company: Join a global leader in the tech industry with a passion for excellence.
  • Benefits: Enjoy hybrid working, eye care, life assurance, and a competitive salary.
  • Other info: Collaborative culture with opportunities for continuous improvement and career growth.
  • Why this job: Make a real impact on system reliability and observability in a dynamic environment.
  • Qualifications: Experience in software engineering and knowledge of Site Reliability Engineering principles.

The predicted salary is between 28800 - 48000 £ per year.

As a Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. You will have software engineering skills, focusing on system reliability and observability. You will monitor the health, performance and availability of critical systems, directly impacting operational efficiency. Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with tools such as Open Telemetry, improve logging practices and develop features for maintainability. You will also help engineer tools and automation for effective service management. Collaboration is key, working across multiple functions to integrate reliability and observability best practices into the software development life cycle. By supporting governance standards set by the central teams, you will foster a culture where these principles are integral to development. Your contributions will ensure our systems meet user demands and enhance overall service performance. This role is eligible for inclusion in the Company’s hybrid working from home policy.

Preferred Skills and Experience

  • Excellent knowledge of Site Reliability Engineering principles, including the creation and management of effective Service Level Indicators (SLI) and Service Level Objectives (SLO) for reliability and customer satisfaction.
  • Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and Pager Duty.
  • Knowledge and experience of modern software development techniques and lifecycles.
  • Experience with Infrastructure as Code (IaC) automation and orchestration tools such as Ansible and Terraform.
  • Prior experience working in a large scale, 24/7 enterprise where system uptime and stability is of paramount importance to the Business.
  • Keen interest of industry trends, particularly Platform Engineering.
  • Proficiency in shell scripting for automation and system management tasks.

What you will be doing

  • Writing and contributing to code that enhances the reliability and observability of services, including telemetry, operational APIs and tooling.
  • Developing and maintaining tools that facilitate effective management of our systems, ensuring they are operationally efficient and resilient.
  • Working with automation and orchestration platforms to automate manual activity and reduce toil.
  • Building sophisticated dashboards using a range of telemetry data and dash boarding technologies like Grafana, Splunk and New Relic.
  • Maintaining and administering existing monitoring and analytic toolsets.
  • Mentoring colleagues in use of new technologies or practices.
  • Actively participating in live incident resolution and post-mortem analysis, providing effective remediation strategies to improve overall system health and prevent future issues.
  • Driving initiatives to enhance system reliability and observability, contributing to a culture of continuous improvement.
  • Collaborating with the central Site Reliability Engineering and Observability teams to establish and uphold standards for reliability and observability, assisting teams in adhering to these practices.
  • Working with IT Operations, providing and supporting the use of critical tooling to enable increasing levels of value to the Business.

Bonus

  • Eye care and Flu Vaccinations
  • Life Assurance

Life at bet365: We are a unique global operator with passion and drive to be the best in the industry. Our values form the foundation of culture and shape the unique way that we work. People are our superpower and we support you to be the best you can be.

Site Reliability Engineer employer: bet365 Group

At bet365 Group, we pride ourselves on being an exceptional employer, offering a dynamic work environment in Stoke-on-Trent that fosters innovation and collaboration. Our hybrid working options provide flexibility, while our commitment to professional development ensures that you have access to in-house training and growth opportunities. Join us to be part of a forward-thinking team dedicated to delivering quality solutions in the exciting world of sports trading.

B

Contact Details:

bet365 Group Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Site Reliability Engineer

Tip Number 1

Network like a pro! Reach out to current Site Reliability Engineers on LinkedIn or at industry events. Ask them about their experiences and any tips they might have for landing a role like this. You never know who might have the inside scoop on job openings!

Tip Number 2

Show off your skills! Create a portfolio showcasing your projects related to system reliability and observability. Include examples of how you've used tools like Grafana or New Relic. This will give potential employers a taste of what you can bring to the table.

Tip Number 3

Prepare for those interviews! Brush up on your knowledge of SLI and SLO principles, and be ready to discuss how you've implemented these in past roles. Practising common interview questions can help you feel more confident when it’s time to shine.

Tip Number 4

Don’t forget to apply through our website! We love seeing applications from passionate candidates who are eager to enhance system reliability. Plus, it’s a great way to ensure your application gets into the right hands quickly.

We think you need these skills to ace Site Reliability Engineer

Site Reliability Engineering principles
Service Level Indicators (SLI)
Service Level Objectives (SLO)
Observability tools (Splunk, New Relic, Grafana, Pager Duty)
Infrastructure as Code (IaC)
Automation and orchestration tools (Ansible, Terraform)
Shell scripting

Some tips for your application 🫡

Tailor Your CV:Make sure your CV reflects the skills and experiences that align with the Site Reliability Engineer role. Highlight your software engineering skills, especially in system reliability and observability, to catch our eye!

Craft a Compelling Cover Letter:Use your cover letter to tell us why you're passionate about enhancing system reliability and how your experience with tools like Grafana or New Relic can make a difference. Show us your personality and enthusiasm!

Showcase Relevant Projects:If you've worked on projects involving Infrastructure as Code or automation tools, be sure to mention them! We love seeing practical examples of your work that demonstrate your expertise in maintaining system uptime.

Apply Through Our Website:We encourage you to apply directly through our website for the best chance of getting noticed. It’s the easiest way for us to keep track of your application and ensure it reaches the right team!

How to prepare for a job interview at bet365 Group

Know Your SRE Principles

Make sure you brush up on Site Reliability Engineering principles, especially around Service Level Indicators (SLIs) and Service Level Objectives (SLOs). Be ready to discuss how you've applied these in past roles or projects, as this will show your understanding of reliability and customer satisfaction.

Familiarise with Observability Tools

Get hands-on experience with tools like Splunk, New Relic, and Grafana before the interview. Being able to talk about how you've used these tools to enhance system observability will set you apart. Maybe even prepare a mini-case study of a project where you implemented these tools effectively.

Showcase Your Automation Skills

Since automation is key in this role, be prepared to discuss your experience with Infrastructure as Code (IaC) tools like Ansible and Terraform. Bring examples of how you've automated processes in previous jobs, and if possible, share any scripts you've written that improved system management.

Collaboration is Key

This role requires working across multiple functions, so be ready to share examples of how you've collaborated with different teams. Highlight any experiences where you helped integrate reliability and observability best practices into the software development life cycle, as this will demonstrate your teamwork skills.