Senior Site Reliability Engineer - Azure Cloud Platform Southampton - Hybrid My client provides advanced SaaS solutions to organisations operating within the public safety and justice sectors. Its technology supports multimedia evidence management and emergency contact centre operations for customers around the world. Due to continued growth, my client is expanding its Cloud Platform Engineering function and is looking for an experienced Senior Site Reliability Engineer. This is a highly hands-on position focused on ensuring that business-critical cloud platforms remain observable, measurable, secure, scalable and reliable. The successful candidate will combine strong Azure engineering expertise with platform automation, observability and operational leadership. Candidates are likely to come from a Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering or Cloud Development background. Security requirement: Applicants must have lived in the UK continuously for the past five years and be eligible to obtain NPPV3 and UK Security Clearance. Role Responsibilities Work as part of the Site Reliability Engineering team responsible for protecting and improving production environments. Manage and prioritise a technical backlog of reliability, scalability and operational improvements. Lead investigations into service outages, performance degradation, platform reliability and cloud expenditure. Conduct detailed root-cause analysis and ensure corrective actions are implemented. Identify repetitive operational activities and replace them with sustainable automation. Provide technical leadership and guidance to Cloud Operations, Support, DevOps and Engineering teams. Establish and maintain service level objectives, service level agreements, service level indicators and error budgets. Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices. Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry. Develop custom application and platform metrics to improve operational visibility. Create advanced queries, dashboards and alerts for distributed microservices. Develop reusable Bicep or Terraform modules for monitoring and cloud infrastructure. Support and improve production Kubernetes environments, particularly Azure Kubernetes Service. Review and optimise platform performance, availability, security and cost. Contribute to cloud architecture, technical scoping and the implementation of scalable platform solutions. Support continuous improvement across deployment, provisioning and operational processes. Help ensure platforms and working practices meet relevant security, governance and compliance requirements. Explore opportunities to use AI-assisted tools to improve automation, troubleshooting and engineering productivity.Essential Skills and Experience At least six years' commercial experience within Site Reliability Engineering or a closely related cloud platform role. Demonstrable experience supporting business-critical cloud platforms and live production services. Strong hands-on knowledge of Microsoft Azure. Production experience with Kubernetes and containerised workloads, ideally using AKS. Extensive experience in platform engineering, cloud provisioning and observability. Strong monitoring, alerting and dashboarding experience using technologies such as: Azure Monitor, Grafana, Prometheus, OpenTelemetry, Elasticsearch Experience creating custom metrics, queries, dashboards and alerts for microservices. Advanced scripting or software development skills using PowerShell, Python, C# or a comparable language. Strong Infrastructure as Code experience using Bicep, ARM or Terraform. Experience using Git or another version-control platform. Good knowledge of Microsoft SQL Server, Elasticsearch and structured data formats including YAML, JSON and XML. Strong understanding of microservices architecture, cloud platforms and containerisation. Experience defining or working with SLOs, SLAs, SLIs and error budgets. Excellent troubleshooting and root-cause analysis skills. Experience designing scalable, secure and maintainable cloud solutions. Strong understanding of cybersecurity principles, governance and compliance. Experience operating across both transformation projects and live-service environments.Desirable Skills and Experience Azure DevOps pipeline experience covering CI/CD and automated deployment. Experience developing reusable infrastructure and monitoring modules. Familiarity with AI-enabled engineering and automation tools. Knowledge of security and compliance frameworks such as: ISO 27001, Cyber Essentials Plus, FedRAMP Experience providing technical leadership across multidisciplinary cloud, engineering and support teams. A background in regulated, public-sector or security-sensitive environments.Spectrum IT Recruitment (South) Limited is acting as an Employment Agency in relation to this vacancy
Lead Site Reliability Engineer in Southampton employer: Spectrum IT Recruitment
Join a leading international managed services provider that values your expertise and offers a fully remote position with the flexibility to maintain a healthy work-life balance. With a strong focus on employee growth, you will have the opportunity to take ownership of high-profile compliance projects while enjoying a supportive work culture that encourages continuous improvement and innovation. The monthly visits to the Milton Keynes office foster collaboration and connection, making this an attractive workplace for those seeking meaningful and rewarding employment.