b Overview /b p As Site Reliability Engineer at IMG, you design, build, and operate resilient platforms underpinning our digital, cloud, and live-broadcast services. You will improve reliability and observability while driving automation and disaster recovery readiness across on-prem and cloud environments. You collaborate with engineering and operations to raise release quality and incident response standards. You play a key role in scaling systems for live, business-critical workloads and supporting high-availability workflows. /p b Responsibilities /b ul li Design, build, and maintain reliable, scalable infrastructure across on-prem and cloud environments /li li Improve availability, latency, and efficiency through reliability engineering practices /li li Enhance observability with monitoring, logging, alerting, dashboards, and service health indicators /li li Define SLIs/SLOs, alerting standards, and runbooks for critical services /li li Automate provisioning, configuration, deployment, and recovery using IaC and scripting /li li Collaborate with software, platform, and broadcast engineering teams to improve resilience /li li Serve as escalation point for production incidents and drive post-incident follow-up /li li Lead root cause analysis and preventive actions for incidents /li li Support high availability, backup, failover, and disaster recovery design and testing /li li Enforce security, access control, patching, and best practices across infrastructure /li li Optimize capacity, cost, and performance; produce technical documentation /li li Support live events and critical operational workflows requiring rapid response and clear communication /li li Contribute to planning for new services, migrations, and platform enhancements with resilience in mind /li li Improve platform reliability, stability, and recovery across IMG services /li li Drive reduced mean time to detect/resolve incidents via observability and automation /li li Promote operational ownership and service standards across environments /li li Strengthen resilience for live client-facing workflows through tested failover approaches /li /ul b Key requirements /b ul li Proven experience as Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar /li li Strong knowledge of Linux /li li Hands-on experience with AWS, Azure, or Google Cloud /li li Experience with Docker and Kubernetes /li li Experience with CI/CD tooling and modern software delivery /li li Hands-on with Infrastructure as Code tools like Terraform or CloudFormation /li li Experience with monitoring, logging, alerting, and observability design /li li Solid understanding of networking, security, architecture, and distributed systems /li li Scripting or programming in Python or Bash /li li Experience in high-availability, live production, or business-critical environments /li li Strong troubleshooting, calm decision-making under pressure, and continuous improvement mindset /li li Excellent communication and collaboration with technical and non-technical stakeholders /li /ul ul li calm under pressure /li li collaboration /li li proactive ownership /li li Linux fundamentals /li li AWS/Azure/Google Cloud /li li Docker /li /ul
Site Reliability Engineer, Studios employer: iMG world
IMG is an exceptional employer, offering a dynamic work environment in the heart of London where innovation and collaboration thrive. Employees benefit from competitive salaries, comprehensive health insurance, and a generous holiday allowance, alongside opportunities for professional growth within a leading global sports marketing agency. With a focus on operational excellence and an inclusive culture, IMG empowers its team to excel in their roles while contributing to exciting projects that shape the future of sports broadcasting.