Site Reliability Engineering Lead in Woking

Site Reliability Engineering Lead in Woking

Woking Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
applied

At a Glance

  • Tasks: Lead a team to ensure system reliability and performance while driving innovative initiatives.
  • Company: Applied, a forward-thinking tech company in Woking, Surrey.
  • Benefits: Competitive salary, visa sponsorship, and a commitment to diversity.
  • Other info: Join a collaborative culture focused on continuous improvement and professional development.
  • Why this job: Make a real impact by enhancing system resilience and operational efficiency.
  • Qualifications: Experience in Site Reliability Engineering or DevOps with strong programming skills.

The predicted salary is between 63000 - 77000 Β£ per year.

Site Reliability Engineering Lead at Applied

About the role As the Site Reliability Engineering Lead at Applied, you will play a pivotal role in ensuring the reliability, availability, and performance of our systems.

You will lead a talented team of engineers to design, implement, and maintain scalable infrastructure that supports our innovative applications.

This position requires a strategic mindset, as you will be responsible for driving initiatives that enhance system resilience and operational efficiency.

Your leadership will be crucial in fostering a culture of collaboration and continuous improvement within the engineering team.

Key facts

Location: Woking, Surrey Engagement: Permanent Salary: Competitive, commensurate with experience Visa: Sponsorship available for the right candidate

What you'll do

  • Lead and mentor a team of Site Reliability Engineers, promoting best practices in system reliability and performance.
  • Design and implement robust monitoring and alerting systems to proactively identify and resolve issues before they impact users.
  • Collaborate with software engineering teams to ensure that new features are designed with reliability and scalability in mind.
  • Develop and maintain infrastructure as code using tools such as Terraform or Cloud Formation to automate deployment processes.
  • Establish and enforce service level objectives (SLOs) and service level indicators (SLIs) to measure system performance and reliability.
  • Conduct post-mortem analyses for incidents, identifying root causes and implementing preventive measures to avoid future occurrences.
  • Optimize system performance through capacity planning and load testing, ensuring that applications can handle increased traffic and usage.
  • Drive initiatives for continuous integration and continuous deployment (CI/CD) to streamline development workflows and reduce time to market.
  • Participate in on-call rotations, providing support for critical incidents and ensuring timely resolution of issues.
  • Foster a culture of collaboration and knowledge sharing within the team, encouraging professional development and skill enhancement.
  • Stay current with industry trends and emerging technologies, evaluating their potential impact on our systems and processes.
  • Advocate for and implement security best practices across all aspects of system design and operation.

Requirements

  • Proven experience in a Site Reliability Engineering or Dev Ops role, with a strong focus on system reliability and performance.
  • Proficiency in programming languages such as Python, Go, or Java, with the ability to write clean, maintainable code.
  • Extensive experience with cloud platforms like AWS, Azure, or Google Cloud, including services related to compute, storage, and networking.
  • Strong understanding of containerization technologies such as Docker and orchestration tools like Kubernetes.
  • Familiarity with monitoring and logging tools such as Prometheus, Grafana, ELK Stack, or similar solutions.
  • Excellent problem-solving skills and the ability to work effectively under pressure in a fast-paced environment.
  • Strong communication skills, with the ability to convey complex technical concepts to non-technical stakeholders.
  • Experience with agile methodologies and a passion for continuous improvement and automation.
  • Nice to have
  • Knowledge of configuration management tools such as Ansible, Puppet, or Chef.
  • Experience with database technologies, including SQL and No SQL databases, and their operational management.
  • Familiarity with security best practices and compliance frameworks relevant to cloud infrastructure.
  • Contributions to open-source projects or active participation in the tech community.

Skills & tools

  • Programming Languages: Python, Go, Java
  • Cloud Platforms: AWS, Azure, Google Cloud
  • Containerization: Docker, Kubernetes
  • Monitoring: Prometheus, Grafana, ELK Stack
  • Infrastructure as Code: Terraform, Cloud Formation
  • Configuration Management: Ansible, Puppet, Chef
  • Practical notes
  • This role is based in Woking, Surrey, and may require occasional travel for team meetings or conferences.
  • The position offers a competitive salary that will be determined based on your experience and qualifications.
  • Applied is committed to fostering a diverse and inclusive workplace; we encourage applications from candidates of all backgrounds.

If you are passionate about building reliable systems and leading a team of engineers to achieve excellence, we would love to hear from you.

#J-18808-Ljbffr

Site Reliability Engineering Lead in Woking employer: applied

Applied in Woking is an excellent employer, offering a dynamic work culture that fosters collaboration and innovation. Employees benefit from comprehensive growth opportunities, including professional development and the chance to lead impactful projects in the electrification sector. With a focus on teamwork and a commitment to employee well-being, Applied provides a rewarding environment for those looking to make a meaningful contribution in their careers.

applied

Contact Details:

applied Recruitment Team

We think you need these skills to ace Site Reliability Engineering Lead in Woking

Site Reliability Engineering
DevOps
System Reliability
Performance Optimisation
Infrastructure as Code
Terraform
CloudFormation