At a Glance
- Tasks: Lead a team to ensure system reliability and performance while driving innovative initiatives.
- Company: Applied, a forward-thinking tech company in Woking, Surrey.
- Benefits: Competitive salary, visa sponsorship, and a commitment to diversity.
- Other info: Join a collaborative culture focused on continuous improvement and professional development.
- Why this job: Make a real impact by enhancing system resilience and operational efficiency.
- Qualifications: Experience in Site Reliability Engineering or DevOps with strong programming skills.
The predicted salary is between 63000 - 77000 Β£ per year.
Site Reliability Engineering Lead at Applied
About the role As the Site Reliability Engineering Lead at Applied, you will play a pivotal role in ensuring the reliability, availability, and performance of our systems.
You will lead a talented team of engineers to design, implement, and maintain scalable infrastructure that supports our innovative applications.
This position requires a strategic mindset, as you will be responsible for driving initiatives that enhance system resilience and operational efficiency.
Your leadership will be crucial in fostering a culture of collaboration and continuous improvement within the engineering team.
Key facts
Location: Woking, Surrey Engagement: Permanent Salary: Competitive, commensurate with experience Visa: Sponsorship available for the right candidate
What you'll do
- Lead and mentor a team of Site Reliability Engineers, promoting best practices in system reliability and performance.
- Design and implement robust monitoring and alerting systems to proactively identify and resolve issues before they impact users.
- Collaborate with software engineering teams to ensure that new features are designed with reliability and scalability in mind.
- Develop and maintain infrastructure as code using tools such as Terraform or Cloud Formation to automate deployment processes.
- Establish and enforce service level objectives (SLOs) and service level indicators (SLIs) to measure system performance and reliability.
- Conduct post-mortem analyses for incidents, identifying root causes and implementing preventive measures to avoid future occurrences.
- Optimize system performance through capacity planning and load testing, ensuring that applications can handle increased traffic and usage.
- Drive initiatives for continuous integration and continuous deployment (CI/CD) to streamline development workflows and reduce time to market.
- Participate in on-call rotations, providing support for critical incidents and ensuring timely resolution of issues.
- Foster a culture of collaboration and knowledge sharing within the team, encouraging professional development and skill enhancement.
- Stay current with industry trends and emerging technologies, evaluating their potential impact on our systems and processes.
- Advocate for and implement security best practices across all aspects of system design and operation.
Requirements
- Proven experience in a Site Reliability Engineering or Dev Ops role, with a strong focus on system reliability and performance.
- Proficiency in programming languages such as Python, Go, or Java, with the ability to write clean, maintainable code.
- Extensive experience with cloud platforms like AWS, Azure, or Google Cloud, including services related to compute, storage, and networking.
- Strong understanding of containerization technologies such as Docker and orchestration tools like Kubernetes.
- Familiarity with monitoring and logging tools such as Prometheus, Grafana, ELK Stack, or similar solutions.
- Excellent problem-solving skills and the ability to work effectively under pressure in a fast-paced environment.
- Strong communication skills, with the ability to convey complex technical concepts to non-technical stakeholders.
- Experience with agile methodologies and a passion for continuous improvement and automation.
- Nice to have
- Knowledge of configuration management tools such as Ansible, Puppet, or Chef.
- Experience with database technologies, including SQL and No SQL databases, and their operational management.
- Familiarity with security best practices and compliance frameworks relevant to cloud infrastructure.
- Contributions to open-source projects or active participation in the tech community.
Skills & tools
- Programming Languages: Python, Go, Java
- Cloud Platforms: AWS, Azure, Google Cloud
- Containerization: Docker, Kubernetes
- Monitoring: Prometheus, Grafana, ELK Stack
- Infrastructure as Code: Terraform, Cloud Formation
- Configuration Management: Ansible, Puppet, Chef
- Practical notes
- This role is based in Woking, Surrey, and may require occasional travel for team meetings or conferences.
- The position offers a competitive salary that will be determined based on your experience and qualifications.
- Applied is committed to fostering a diverse and inclusive workplace; we encourage applications from candidates of all backgrounds.
If you are passionate about building reliable systems and leading a team of engineers to achieve excellence, we would love to hear from you.
#J-18808-Ljbffr
Site Reliability Engineering Lead in Woking employer: applied
Applied in Woking is an excellent employer, offering a dynamic work culture that fosters collaboration and innovation. Employees benefit from comprehensive growth opportunities, including professional development and the chance to lead impactful projects in the electrification sector. With a focus on teamwork and a commitment to employee well-being, Applied provides a rewarding environment for those looking to make a meaningful contribution in their careers.