Site Reliability Engineer

Site Reliability Engineer

Full-Time 63000 - 77000 Β£ / year (est.) Home office (partial)
N

At a Glance

  • Tasks: Transform the SDLC environment with automation and ensure system reliability.
  • Company: Join Natobotics, a forward-thinking tech company in London.
  • Benefits: Enjoy hybrid work, competitive pay, and opportunities for professional growth.
  • Other info: Collaborative culture with a focus on continuous improvement and innovation.
  • Why this job: Make a real impact by enhancing system performance and reliability.
  • Qualifications: 15+ years of experience in SRE, automation, and cloud technologies.

The predicted salary is between 63000 - 77000 Β£ per year.

Overview

Join to apply for the Site Reliability Engineer role at Natobotics.

Experience Level: 15+ Years.

A Site Reliability Engineer is responsible for transforming the SDLC environment with engineering-focused role that emphasizes system reliability, automation, and performance in a non-production setting.

Responsibilities

  • Automate environment lifecycle: Develop Infrastructure as Code (Ia C) to automate provisioning, teardown, and configuration of test environments, integrating them with the CI/CD pipeline.
  • Establish service level objectives (SLOs): Define and measure SLIs for test environments, such as availability and provisioning time.
  • Monitor environment health and performance: Use observability tools like Prometheus and Grafana to track the health of test environments, identify bottlenecks, and resolve issues proactively, not reactively.
  • Manage incident response: Lead the incident management process for test environment issues, conducting blameless post-mortems to understand the root causes and implement lasting fixes.
  • Minimize toil: Automate manual, repetitive tasks associated with test environments to free up engineering time for more strategic work.
  • Strategic and cultural responsibilities
  • Drive continuous improvement: Analyze environment performance data, incident reports, and post-mortems to identify opportunities for continuous improvement and innovation.
  • Balance reliability and speed: Use an "error budget" for test environments.

If environments are highly reliable, teams can use the budget for quicker feature development.

If reliability is low, the focus shifts to improving stability.

  • Instil a reliability culture: Promote a blameless culture around test environment incidents, encouraging shared ownership and collaboration between development, QA, and SRE teams.
  • Capacity planning: Anticipate the future resource needs of test environments by analysing usage patterns and project forecasts.

Ensure the infrastructure can scale to meet demand.

  • Advance test data management: Work with Test Data Managers to ensure that test data is not only readily available but also consistent, compliant, and automatically provisioned with the environments.
  • Technical Skills
  • Expertise in tooling: Proficiency with monitoring and logging tools (e. g., Prometheus, Splunk, Grafana), CI/CD platforms (e. g., Jenkins, Git Lab CI), and configuration management tools (e. g., Ansible, Terraform).
  • Cloud infrastructure knowledge: Deep understanding of cloud platforms like AWS, including experience with containerization technologies (Docker, Kubernetes) and serverless computing.
  • Scripting and programming: Strong scripting skills in languages such as Python or Bash to automate environment management tasks.
  • Systems and networking knowledge: Solid understanding of Linux systems, networking concepts, and database management.
  • Soft Skills
  • Leadership and influence: The ability to champion SRE practices and influence technical and business stakeholders across different teams.
  • Problem-solving: Strong analytical and debugging skills for investigating and resolving complex environment issues under pressure.
  • Communication: Excellent communication and collaboration skills to bridge the gap between development, QA, and operations teams.
  • Adaptability: A proactive and adaptable mindset to keep pace with evolving technology and development methodologies.
  • Employment and Location
  • Seniority level: Mid-Senior level

Note: Referrals increase your chances of interviewing at Natobotics by 2x.

#J-18808-Ljbffr

Site Reliability Engineer employer: Natobotics

N Consulting Ltd is an exceptional employer that values innovation and collaboration, offering a dynamic work culture where experienced professionals can thrive. With a focus on employee growth, you will have the opportunity to work on cutting-edge projects in a remote setting, allowing for flexibility while contributing to impactful data solutions. Join us to be part of a team that prioritises quality, compliance, and the advancement of your career in the exciting field of data engineering.

N

Contact Details:

Natobotics Recruitment Team

We think you need these skills to ace Site Reliability Engineer

SQL
Python
Automation
Communication Skills
Data Governance
Problem-Solving Skills
Data Engineering