Operations Engineer

Operations Engineer

Full-Time 40500 - 49500 £ / year (est.) Home office (partial)
AGS

At a Glance

  • Tasks: Support and optimise critical technology services, ensuring stability and performance.
  • Company: Join a forward-thinking company with a focus on innovation and collaboration.
  • Benefits: Enjoy a competitive salary, flexible hybrid work, and great career development opportunities.
  • Other info: Work in a dynamic environment with opportunities for continuous improvement and growth.
  • Why this job: Make a real impact by enhancing the reliability of essential digital services.
  • Qualifications: Experience in IT operations and strong troubleshooting skills are essential.

The predicted salary is between 40500 - 49500 £ per year.

Hybrid - 1 day in the office every 1 - 2 months

The Opportunity

Our client is looking for an Operations Engineer to help maintain the stability, availability and performance of a range of critical technology services and platforms.

This is a key operational position, supporting the technology estate behind important digital and commercial services.

The role will cover enterprise applications, integrations, customer-facing systems, platforms and associated infrastructure, with a strong emphasis on proactive monitoring, effective incident resolution and continual service improvement.

You will work across a broad technology environment, identifying potential issues before they impact users, responding to operational incidents and helping ensure that services remain reliable, secure and capable of supporting business requirements.

The position involves close collaboration with teams across Engineering, Platform, Integration, Product, Cybersecurity, Enterprise IT and Service Management.

You will contribute to areas including incident response, change and release activities, operational governance, platform optimisation and service continuity.

Key Responsibilities

  • Provide day-to-day operational and technical support across enterprise applications, platforms, integrations and related technologies.
  • Monitor the health, availability and performance of critical services through appropriate monitoring, alerting and observability solutions.
  • Investigate technical incidents, diagnose underlying problems and restore affected services as efficiently as possible.
  • Undertake root cause analysis and work towards sustainable fixes that reduce the likelihood of repeat incidents.
  • Assist with the coordination and management of major incidents, working with relevant internal teams and third-party suppliers where appropriate.
  • Support production releases, deployments and associated checks to confirm that environments are ready and operationally stable.
  • Partner with Engineering and Platform teams to ensure new technology, functionality and releases can be effectively supported once introduced into production.
  • Look for ways to improve and automate operational activities, including maintenance, reporting, alerting and routine processes.
  • Ensure operational procedures and activities comply with relevant governance, cybersecurity, compliance and change control standards.
  • Produce and maintain clear technical documentation, operational procedures, support information and runbooks.
  • Contribute to continuous improvement programmes designed to strengthen reliability, platform stability and overall operational efficiency.
  • Build effective working relationships with Product, Engineering, Platform, Integration and Service Management teams to ensure high-quality operational delivery.

The successful candidate will bring

  • Demonstrable experience providing operational support for enterprise applications, digital platforms or cloud-based technology.
  • Strong technical investigation and troubleshooting skills, supported by a logical and analytical approach to problem solving.
  • Previous experience within IT Operations, Support Engineering or Service Management environments.
  • Experience working with production systems and business-critical services where availability and reliability are important.
  • A good understanding of monitoring, logging, alerting and wider observability practices.
  • Practical experience working with incident, problem, change and release management processes.
  • Strong communication and interpersonal skills, with the ability to work effectively with a variety of technical teams and business stakeholders.

The following would be advantageous

  • Background within automotive, manufacturing, retail or other large-scale enterprise technology environments.
  • Familiarity with cloud technologies, CI/CD platforms, Dev Ops methodologies and infrastructure automation.
  • Experience supporting Salesforce, APIs, system integrations, digital services or enterprise Saa S platforms.
  • Knowledge of Site Reliability Engineering (SRE), operational engineering methodologies or other reliability-focused approaches.
  • #J-18808-Ljbffr

Operations Engineer employer: AGS

Join a dynamic and innovative team in London as a Data Engineer, where you'll have the chance to shape the future of our data platform while enjoying a hybrid work model. We pride ourselves on fostering a collaborative work culture that encourages professional growth and offers opportunities to take ownership of impactful projects. With a focus on data governance and cutting-edge technologies like Databricks, this role not only promises meaningful work but also the chance to be part of a forward-thinking organisation committed to excellence.

AGS

Contact Details:

AGS Recruitment Team

We think you need these skills to ace Operations Engineer

Operational Support
Technical Investigation
Troubleshooting Skills
Analytical Approach
Incident Management
Problem Management
Change Management