At a Glance
- Tasks: Build and improve a highly available cloud platform for mission-critical services.
- Company: Innovative tech company focused on cloud technologies and collaboration.
- Benefits: Flexible remote work, professional development, and modern Apple equipment.
- Other info: Inclusive workplace encouraging diverse applicants and continuous improvement.
- Why this job: Join a dynamic team and make a real impact on cloud reliability.
- Qualifications: Experience with Kubernetes, AWS, and strong scripting skills required.
The predicted salary is between 60000 - 80000 £ per year.
Role Overview
B2B Contract
We're looking for a Senior Site Reliability Engineer to help build, operate, and continuously improve a highly available cloud platform supporting mission-critical production services.
In this role, you'll work at the intersection of cloud infrastructure, software engineering, and operations, helping engineering teams build reliable, scalable systems while driving automation, observability, and operational excellence.
You'll play a key role in strengthening platform reliability, improving incident response, and embedding SRE best practices throughout the software development lifecycle.
Key Responsibilities
- Maintain the reliability, availability, and performance of production and pre‑production environments.
- Monitor platform health and improve alerting, automation, and operational processes.
- Respond to production incidents, participate in root cause analysis, and implement long‑term improvements.
- Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards.
- Partner with software engineers to improve application reliability throughout the development lifecycle.
- Develop and maintain operational documentation, troubleshooting guides, and runbooks.
- Automate repetitive operational tasks to improve efficiency and reduce manual intervention.
- Participate in on‑call rotations while continuously improving incident response processes.
- Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams.
Requirements
- Bachelor's or Master's degree in Engineering, Computer Science, or a related field.
- Strong experience operating Kubernetes or other container orchestration platforms.
- Experience supporting large‑scale production services.
- Hands‑on experience with AWS.
- Experience with Prometheus, Grafana, and ELK.
- Strong scripting skills (Bash, Python, or Go).
- Experience administering Linux‑based production environments.
- Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible.
- Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing).
- Excellent troubleshooting, communication, and collaboration skills.
- A proactive mindset with a passion for automation and reliability.
- Nice to Have
- Experience with SIP or Vo IP technologies.
- Familiarity with My SQL or Postgre SQL.
- Experience with Redis or other No SQL databases.
- What's on Offer
- Long‑term, full‑time collaboration.
- Flexible remote working environment.
- Professional development opportunities, including training and technical learning.
- The opportunity to work on innovative cloud technologies used by customers worldwide.
- Collaborative engineering culture focused on knowledge sharing and continuous improvement.
- Modern Apple equipment provided.
- Diversity and Inclusion Commitment
We are dedicated to creating and sustaining an inclusive, respectful workplace for all—regardless of gender, ethnicity, or background.
We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast‑paced, supportive team.
#J-18808-Ljbffr
Senior Site Reliability Engineer (Cloud Platform) in Manchester employer: Salve.Lab
Salve.Lab is an exceptional employer that fosters a dynamic and inclusive work culture, offering employees the chance to thrive in a rapidly growing HR Technology environment. With a focus on professional development and collaboration, team members are encouraged to innovate and contribute to meaningful projects while enjoying the flexibility of a hybrid work model in vibrant London. The company values diversity and provides unique opportunities for growth, making it an attractive place for motivated individuals looking to make a significant impact in their careers.