Senior Infrastructure Engineer (Linux & Cloud Automation) in Manchester

Senior Infrastructure Engineer (Linux & Cloud Automation) in Manchester

Manchester Full-Time On-site
S

Job Details: Senior Infrastructure Engineer (Linux & Cloud Automation)

Full details of the job.

Vacancy Name

Vacancy No

Vacancy No VN970

Function

Function Syn - Delivery & Operations

Work Location

Work Location One St Peter's Square, Manchester, M2 3DE

Basis

Basis Permanent

Full Time/Part Time

Full Time/Part Time Full time

Employment Duration

Employment Duration -

Hours Per Week 35.00

Drivers Licence Required

Drivers Licence Required No

Benefits Competitive Salary and Benefits package will be provided to the successful candidate.

About the Role

  • Minimum of 5–7 years of experience in enterprise infrastructure engineering, ideally within managed services or a multi-customer MSP environment.
  • Demonstrated ownership of large-scale Linux patching programmes using Ansible, including scheduling, deployment, validation, remediation, and compliance reporting.
  • Hands‑on experience with Ansible — playbook development and execution for patching, configuration management, and compliance enforcement.
  • Proficiency in Microsoft Azure — specifically Linux VM management, Azure Update Manager, Azure Backup, Recovery Services Vaults, NSG/UDR configuration, and Azure networking.
  • Strong Linux administration skills across RHEL, CentOS, Ubuntu, and Oracle Linux — package management, service management, log analysis, and security hardening.
  • Experience with Infrastructure‑as‑Code using Bicep and/or ARM templates for repeatable Azure deployments.
  • Familiarity with Azure DevOps — CI/CD pipelines, repos, and release management for infrastructure automation.
  • Willingness to participate in a 1-in-5 weekly on‑call rotation providing 24×7 P1 incident response.

Position Overview

As a Senior Infrastructure Engineer (Linux & Cloud Automation) at Synapse360, you are responsible for the operational delivery and automation of Linux‑based infrastructure across managed service customers. You will own the end‑to‑end patching lifecycle for Linux server estates using Ansible, and manage Azure Update Manager policies for Linux VMs across Azure. You will manage and remediate backup operations for Linux workloads via Azure Backup, implement infrastructure‑as‑code deployments using Bicep and ARM templates, maintain CI/CD pipelines via Azure DevOps, and handle escalations from the Infrastructure team. This is a demanding, hands‑on engineering role requiring strong proficiency in Linux administration, automation tooling, Microsoft Azure, and DevOps practices. You will participate in the weekly on‑call rota and must be capable of independently resolving P2 incidents and supporting the Infrastructure Lead during P1 major incident response.

What we expect

  • Minimum of 5–7 years of experience in enterprise infrastructure engineering, ideally within managed services or a multi‑customer MSP environment.
  • Demonstrated ownership of large‑scale Linux patching programmes using Ansible, including scheduling, deployment, validation, remediation, and compliance reporting.
  • Hands‑on experience with Ansible — playbook development and execution for patching, configuration management, and compliance enforcement.
  • Proficiency in Microsoft Azure — specifically Linux VM management, Azure Update Manager, Azure Backup, Recovery Services Vaults, NSG/UDR configuration, and Azure networking.
  • Strong Linux administration skills across RHEL, CentOS, Ubuntu, and Oracle Linux — package management, service management, log analysis, and security hardening.
  • Experience with Infrastructure‑as‑Code using Bicep and/or ARM templates for repeatable Azure deployments.
  • Familiarity with Azure DevOps — CI/CD pipelines, repos, and release management for infrastructure automation.
  • Willingness to participate in a 1-in-5 weekly on‑call rotation providing 24×7 P1 incident response.

Areas of Responsibility

  • Patching: Own the monthly patching cycle for Linux servers. Deploy patches via Ansible playbooks. Manage patch scheduling, pre‑patch snapshots, deployment waves, post‑patch validation, and remediation. Achieve >95% patching compliance SLA. Manage Azure Update Manager policies for the Linux Azure estate — configure maintenance windows, deploy OS‑level patches, validate compliance, and remediate failures.
  • Backup Management: Manage Azure Backup policies for Linux workloads, monitor Recovery Services Vaults, perform restore testing, and maintain backup compliance. Develop and maintain scripted backup validation routines. Target >98% backup success rate.
  • Infrastructure‑as‑Code & Automation: Develop and maintain Bicep and ARM templates for Azure resource deployments. Build and manage Azure DevOps CI/CD pipelines for infrastructure provisioning and configuration. Maintain Ansible playbooks for configuration management, compliance enforcement, and operational automation.
  • VM Lifecycle Management: Provision, configure, snapshot, resize, and decommission Linux virtual machines on Azure (IaaS). Manage VM images and ensure compliance with customer standards.
  • Azure Key Vault & Secrets Management: Manage secrets, certificates, and keys within Azure Key Vault. Implement certificate rotation and integrate Key Vault with automated deployments.
  • Change Management: Prepare and implement standard and normal RFCs. Document changes fully including risk assessment, rollback procedures, and post‑implementation reviews. Attend weekly CAB as required.
  • T1 Escalation Handling: Receive and resolve Linux and cloud automation incidents escalated by T1 engineers. Provide guidance and knowledge transfer to T1 team members.
  • ServiceNow & Reporting: Maintain accurate ticket records, contribute to SLA reporting, and ensure all work is logged against the correct customer and category.

On-Call Commitment

This role participates in a shared five‑person weekly on‑call rotation. Each engineer is primary on‑call for one week in every five. During on‑call periods, you are expected to respond to P1 critical alerts within 15 minutes (24×7), independently resolve P2 incidents, and support the Infrastructure Lead during major incident response. On‑call compensation is provided in line with Synapse360's standard on‑call policy.

Ideal Candidate Characteristics

  • Patching Ownership: You are the patching authority for Linux estates. Full lifecycle ownership — scheduling, risk assessment, Ansible playbook deployment, validation, remediation, and compliance reporting. Patching is a critical SLA metric (>95% compliance) and must be treated as a controlled change.
  • Automation & DevOps Mindset: You must be comfortable building and maintaining Ansible playbooks, Bicep templates, and Azure DevOps pipelines. The role demands a strong drive toward automation, repeatability, and infrastructure‑as‑code principles.
  • Backup Operations: You own the remediation of Azure Backup failures for Linux workloads and scripted backup validation. You must understand backup policies and Recovery Services Vault configurations to troubleshoot and maintain compliance.
  • Change Management Discipline: You will prepare and implement standard and normal RFCs, ensuring full documentation, risk assessment, rollback plans, and post‑implementation validation. You will attend and present at the weekly CAB as required.
  • Escalation Handling: You are the escalation point for T1 engineers on Linux and cloud automation issues. You must be able to receive a partially triaged incident and drive it to resolution without unnecessary re‑escalation to Infrastructure Lead.

Important Attributes

  • Patching Ownership: You are the patching authority for Linux estates. Full lifecycle ownership — scheduling, risk assessment, Ansible playbook deployment, validation, remediation, and compliance reporting. Patching is a critical SLA metric (>95% compliance) and must be treated as a controlled change.
  • Automation & DevOps Mindset: You must be comfortable building and maintaining Ansible playbooks, Bicep templates, and Azure DevOps pipelines. The role demands a strong drive toward automation, repeatability, and infrastructure‑as‑code principles.
  • Backup Operations: You own the remediation of Azure Backup failures for Linux workloads and scripted backup validation. You must understand backup policies and Recovery Services Vault configurations to troubleshoot and maintain compliance.
  • Change Management Discipline: You will prepare and implement standard and normal RFCs, ensuring full documentation, risk assessment, rollback plans, and post‑implementation validation. You will attend and present at the weekly CAB as required.
  • Escalation Handling: You are the escalation point for T1 engineers on Linux and cloud automation issues. You must be able to receive a partially triaged incident and drive it to resolution without unnecessary re‑escalation to Infrastructure Lead.

Areas of Responsibility

  • Monitoring & Alert Management: Manage and triage alerts from LogicMonitor and Azure Monitor across environments. Correlate alerts, close false positives, and elevate genuine incidents. Maintain monitoring dashboards and ensure alert thresholds remain appropriate.
  • ServiceNow Ticket Management: Own the ticket lifecycle for incoming incidents and service requests. Triage, categorise, prioritise, and either resolve or escal.
  • Backup Validation: Perform daily Veeam Backup & Replication job status checks (~1,290 servers) and Azure Backup status validation. Log failures, attempt basic remediation (re‑run failed jobs), and escal persistent failures to Senior Engineers.
  • Patching Support: Assist Senior Engineers during monthly patch cycles. This includes pre‑patch checks, server reboots, and post‑patch validation across WSUS/SCCM (Windows) and Ansible‑managed (Linux) estates. Also assist with Azure Update Manager deployment validation and failed‑patch reporting.
  • VM Troubleshooting: Perform basic troubleshooting on Azure VMs (connectivity, disk, performance) and vSphere VMs (console access, snapshot management, resource contention). Independently resolve P3/P4 VM issues.
  • Windows Server Administration: Basic administration including service restarts, event log analysis, disk space management, user access troubleshooting, and DNS/DHCP validation.
  • Incident First Response (On‑Call): During on‑call periods, act as the first responder for all P1 alerts. Perform initial triage, engage vendor support if needed, communicate status

#J-18808-Ljbffr

Senior Infrastructure Engineer (Linux & Cloud Automation) in Manchester employer: Salesforce Sites

Build-A-Bear is an exceptional employer that fosters a vibrant and inclusive work culture in the heart of Central London. With a strong emphasis on employee growth, we provide comprehensive training and development opportunities, ensuring our team members thrive in their roles while delivering outstanding guest experiences. Join us to be part of a dynamic environment where creativity meets dependability, and every day brings new opportunities to inspire and connect with others.

S

Contact Details:

Salesforce Sites Recruitment Team