IT & Service Delivery Engineer

IT & Service Delivery Engineer

Full-Time No working from home possible
O

Business Unit

VBU - Trakm8

Job Description

IT & SERVICE DELIVERY ENGINEER

Department:

R & D

Reporting to:

GROUP IT & SERVICE DELIVERY MANAGER

Job Purpose

Based at Trakm8's head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8's telematics and optimisation platforms and of the corporate IT services. The role spans a mixed environment of corporate office, on-premises data centres and AWS. Customer platforms process data in real time, alongside the standard business IT services used by employees. Reliable delivery of both is critical.

Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery. Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.

Key Responsibilities:

Service Availability & Incident Management

  • Monitor the health, capacity and performance of the telematics and optimisation platforms, corporate IT services and supporting infrastructure, taking action to maintain agreed levels of availability and throughput.
  • Investigate and resolve incidents across Linux and Windows servers, databases, networks, storage, container platforms, end user devices and application deployments, owning each from first report or alert through to root cause and permanent fix.
  • Carry out structured root cause analysis for recurring or significant issues, implementing corrective and preventative actions.
  • Communicate incident progress, risks and resolution clearly to technical and non-technical stakeholders.

Infrastructure Engineering & Delivery

  • Design, build, configure and maintain secure, resilient server infrastructure across data centre, virtualised and AWS environments.
  • Manage and improve containerised services, networking, storage, databases and load balancing.
  • Own backup, replication and restore testing, and contribute to capacity planning and disaster recovery, keeping recovery points and recovery times fit for purpose and evidenced.
  • Work with development, platform and service teams to support reliable application deployment and end-to-end service performance.

Windows, Directory & Endpoint Services

  • Administer Active Directory, Group Policy, DNS and internal certificate services across the domain.
  • Own server and workstation patching and endpoint protection coverage, including approvals, safeguard holds, deployment rings.
  • Administer accounts, access rights and licensing, and maintain inventory, imaging and software deployment tooling.
  • Respond directly to user requests and access issues, taking ownership through to resolution within the team’s operating model.
  • Maintain internal IT services including LAN, wireless, VOIP and mobile telephony, and support IT procurement, supplier management and hardware refresh.

Automation, Monitoring & Continuous Improvement

  • Develop and maintain automation and scripts for repeatable operational tasks.
  • Maintain effective monitoring, alerting and dashboards across metrics and log platforms, reviewing thresholds and coverage to identify issues early and reduce avoidable incidents.
  • Build self-healing and auto-remediation where it is safe to do so, for example automated rebalancing of workloads in response to queue lag or resource pressure.
  • Identify technical debt and resilience gaps and recommend proportionate improvements, making effective use of modern tooling including AI-assisted development and automation.

Security, Documentation & Collaboration

  • Operate infrastructure with a security-first approach, applying access controls, hardening and vulnerability remediation in line with company policies.
  • Support logging, monitoring and security tooling, and assist with investigation and remediation of security alerts.
  • Plan and deliver changes and project work through the RFC and release process, booking into agreed release windows and providing test, rollback and post-implementation evidence.
  • Produce recurring operational and service reporting covering availability, capacity, patching, backup and security, and maintain accurate asset, licence and configuration records.
  • Create and maintain clear technical documentation, runbooks and support procedures for new and existing solutions.

Technology Environment

The environment currently includes the following technologies. The list provides context for the role and is not intended to mean that experience in every technology is essential. The expectation is competence in several of these groups and the ability to pick up the rest.

Service delivery platform

  • Oracle Linux and Windows operating systems across VMware vSphere, Proxmox and AWS
  • Kubernetes and Docker, using MetalLB and Kong API gateway
  • MySQL, Cassandra (Scylla), CockroachDB, Redis, Kafka, Elasticsearch, PostgreSQL and RabbitMQ
  • Ansible and AWX, with scripting using Bash, Python or Perl
  • TCP/IP networking, DNS, routing, VPNs, BGP, SMTP, HTTP and HTTPS
  • Palo Alto firewalls; HAProxy, NGINX and dynamic DNS services
  • Grafana, Prometheus, Graphite (ClickHouse), InfluxDB and Icinga, with alerting into OpsGenie
  • ELK for application logging and Wazuh for platform host security monitoring
  • Git, Bitbucket, Jira and Bamboo CI/CD pipelines
  • iSCSI SAN, storage pools, volumes and snapshots
  • AWS services including EC2, Route53, RDS, ELB, VPC, Multi-AZ subnets, VPN, Transit Gateway and Customer Gateway

Corporate IT

  • Windows Server and Windows desktop estates, Active Directory, Group Policy, DNS, DHCP, DFS and file and print services with Linux (Ubuntu)
  • Microsoft 365 and Entra ID, Intune, including Exchange Online, SharePoint, Teams, multi-factor authentication and Conditional Access
  • VMware vSphere and vCenter, with Veeam Backup & Replication and object storage for offsite copies
  • Endpoint management, patching and software deployment tooling, with BitLocker disk encryption
  • CrowdStrike Falcon endpoint protection and UTMStack SIEM
  • Palo Alto firewalls managed through Panorama, with GlobalProtect remote access
  • LAN, wireless, VOIP and mobile telephony across multiple UK sites
  • Puppet for configuration management and PowerShell for scripting, with Icinga and OpsGenie alerting shared across both estates
  • Internal certificate services, TLS and PKI, split-horizon DNS, LDAP and Kerberos
  • MS SQL behind business systems, plus the Atlassian suite, service desk, intranet and digital signage platforms

Essential Skills / Experience / Competency Required:

  • Strong problem solving, with the ability to pick up an unfamiliar system quickly, retain it, and carry what you learn across into unrelated parts of the estate. Recognising that a fault in one area has the same shape as one you solved somewhere else entirely is worth more here than depth in any single technology.
  • Hands-on experience administering Linux in a production, business-critical environment, including fault diagnosis, performance management, patching and security hardening.
  • Solid Windows Server administration, including Active Directory and Group Policy.
  • Experience administering Microsoft 365 and Entra ID.
  • A working understanding of TCP/IP networking, DNS, routing, VPNs and firewall policy: enough to establish whether a fault sits in the application, the host, or a rule upstream.
  • A working understanding of Kubernetes and containerised services: enough to operate and diagnose them day to day. Deeper experience is welcome but not essential.
  • Scripting and automation in at least one of Bash, Python or PowerShell, and a preference for putting configuration into version control rather than doing things by hand.
  • Methodical, evidence-led troubleshooting: working from log, packet and configuration evidence rather than assumption, including faults that do not originate where they first appear.
  • Clear written and verbal communication, with the ability to explain technical issues and risks to different audiences.
  • A full UK driving licence is required, as occasional travel to data centres or other locations may be required to support IT and platform services.
  • Participation in the team's 24x7 on-call rota, with flexible working outside normal hours where reasonably required to protect or restore critical services.

Desirable Skills / Experience:

Platform and cloud

  • Depth in Kubernetes administration, cluster operations and container networking.
  • AWS, VMware or other virtualisation and cloud platforms.
  • Databases in production, such as MySQL, PostgreSQL, MS SQL, Cassandra, Kafka or Elasticsearch.
  • API gateway technologies such as Kong and measuring or reporting API performance.
  • Telemetry, IoT, high-throughput data platforms or similarly critical real-time services.
  • Storage, SAN and data centre experience covering physical security, networking and compute.

Corporate IT

  • Endpoint management, software deployment and patching tooling such as PDQ, WSUS or Intune.
  • Backup and recovery tooling such as Veeam, including restore testing.
  • Providing IT support to end users directly, with the confidence to deal with colleagues at any level.

Automation and monitoring

  • Configuration management with Ansible, Puppet or AWX.
  • Monitoring, alerting and dashboarding with Icinga, Grafana, Prometheus or similar.
  • Infrastructure-as-code, CI/CD tooling or automated testing for infrastructure changes.
  • AI-assisted coding and automation tooling used to accelerate operational and scripting work.

Security and assurance

  • Firewall policy administration, ideally Palo Alto, and load balancing or reverse proxy work.
  • Endpoint protection, vulnerability management or SIEM platforms.
  • System build following NIST standards, hardening, SELinux and server firewalls.
  • Information security frameworks such as ISO 27001 or Cyber Essentials.

Ways of working

  • Working within a formal RFC and change window process, including rollback planning and evidence.
  • Moving live services off unsupported platforms and end-of-life operating systems without an outage.
  • Disaster recovery, capacity planning and service continuity testing.

Personal Attributes:

  • Calm, methodical and solutions-focused, particularly when responding to live service issues.
  • Takes ownership and follows actions through to resolution.
  • Proactive in identifying risks, improvement opportunities and emerging capacity or resilience concerns.
  • Security-conscious and careful when working with critical systems and customer data.
  • Collaborative and approachable, with a willingness to share knowledge and support colleagues.
  • Organised and adaptable, able to balance planned work with changing operational priorities.
  • Committed to clear documentation, continuous learning and improving how the team works.

Our Values:

  • We’re a close-knit team who care about each other, our customers and making a positive impact, our values drive who we are and the decisions we make:
  • Caring & Committed – We show up for each other and deliver with dedication
  • Make a Difference – We create impact that matters
  • Relentless – We push forward until we succeed
  • Do the Right Thing – We choose integrity every time
  • Joy of the Journey – We enjoy the ride and celebrate along the way

Health & Safety:

As a Trakm8 Group employee, you have the following H&S responsibilities:

  • To comply with the Group H&S policy and procedures
  • To take care of your own health and safety and that of people who may be affected by what you do (or what you do not do)
  • To co-operate with others on health and safety
  • To not interfere with, or misuse, anything provided for your health, safety or welfare
  • To wear/use the correct Protective Equipment at all times
  • To follow the training you have received at all times

Information Security:

As a Trakm8 Group employee, you have the following Information Security responsibilities:

  • To follow the Group Information Security policies at all times
  • To follow the Trakm8 Secure Development Policy at all times;
  • To ensure that passwords under your responsibility are kept confidential
  • To keep Group or Company information secure and confidential at all times
  • To inform your manager immediately if you detect, suspect or witness an incident that may be a breach of security

This role requires screening in the below areas:

  • References
  • Pre-Employment Credit Check
  • Basic Disclosure

The above job description is meant to describe the general nature and level of work being performed; it is not intended to be construed as an exhaustive list of all responsibilities, duties and skills required for the position. All job descriptions are subject to possible modifications in line with the needs of the business.

IT & Service Delivery Engineer employer: Omegro

Omegro is an exceptional employer that fosters a collaborative and innovative work culture, perfect for those passionate about cutting-edge technology in the automotive sector. With a strong emphasis on employee growth, you will have opportunities to mentor others while working on impactful projects that blend embedded systems with AI. Located in a vibrant tech hub, our team enjoys a supportive environment that encourages creativity and professional development.

O

Contact Details:

Omegro Recruitment Team