Overview A subsidiary of Publicis Groupe, Epsilon is a leading provider of multi-channel marketing services, technologies, and database solutions. We do more than collect and store data, and we might be the most important Internet company you've never heard of. Join our team for your chance to work in the digital marketing space and solve meaningful problems on a massive scale-and have fun doing it. The System and Platform Operations Director is a technical leadership role that is responsible for the support, reliability and stability of Epsilon Retail Media production systems, environments and offerings. The team owns the reliability vision for the company, driving continuous improvement through a combination of development and operations initiatives as well as process excellence. This position and their team has solid-line responsibility for operations including the deployment, management, monitoring, reporting, troubleshooting, and repair of production systems. Core to the success of the role is to provide a premium customer support experience focused on a "center of excellence" that allows for a full-service delivery support cycle. This role is responsible for managing the Platform Operation Team centralized within a single geo-region, orchestrating the regional teamwork, serving with both technical and professional support, and championing the company values. The Platform Operations Engineer works closely with the Engineering team to ensure ongoing system stability and supports the Technical Account Managers from an environment's perspective. The Platform Operations team is responsible for supporting all retailers once they are live. Critically important is how this team collaborates and liaises with other teams such as Customer Support, Technical Account Management, Engineering and Customer Success teams. What you'll do
- Operational Practices
- + Establish and manage operational practices and ensure we design, implement and operate a support model that is fit for purpose for our future.
- + Implement proactive solutions for incident and problem detection, response and remediation and continuous improvement
- + Owner of the operational integrity of all production environments.
- Production Monitoring and Operational Reporting
- + Adopt a "Measure Everything" approach to ensure that internal service level objectives and customer service levels agreements are exceeded including executive level reporting on operational health metrics such as SLAs, incident resolution, performance, availability, reliability, capacity etc.
- Customer Support & Incident Management
- + Own incident management processes and on call response.
- + Take ownership of complex issues related to performance, reliability, and scalability and leading resolution of serious incidents and events including communications with customers and wider stakeholders.
- Change Management
- + Uphold processes and procedures to manage change across production platforms
- + Provide insight and expertise on how customers will perceive the changes or impacts to customers to drive customer organization change management and communication.
- + Empower the Delivery teams to release new products, features, updates and fixes quickly, while ensuring Platforms remain reliable and stable.
- System Reliability
- + Work with the wider Engineering, Product, Delivery and Security teams to ensure that appropriate attention is given to production/system reliability.
- + Establish Operational Practices in conjunction with the Product and Engineering teams (e.g. understanding how product feature development could affect the system's overall reliability and performance).
- + Provide delivery status information on System Reliability initiatives to the IT Leadership Team and additional stakeholders with a focus and ensure proper communication concerning changes to agreed milestones or challenges, risks and blockers that may affect the outcome or agreed completion dates (with proactive suggestions to resolve)
- IT Service Management
- + Execute Service Management processes including Change, Config, Service Level, Performance, Incident and Problem Management to deliver a high level of support and system availability
- + Leverage industry standards and best practices for
#J-18808-Ljbffr
Manager, System and Platform Operations employer: Publicis Groupe
Starcom is an exceptional employer that prioritises its people, fostering a welcoming and supportive culture where every individual is encouraged to reach their fullest potential. With a strong focus on diversity and inclusion, alongside comprehensive benefits such as enhanced parental leave and reflection days, employees are empowered to thrive both personally and professionally. Located in the UK, this role offers the chance to work with leading global brands while being part of a high-performing team dedicated to innovation and excellence in media investment.