At a Glance
- Tasks: Lead a team to enhance infrastructure reliability and implement engineering best practices.
- Company: Join a leading financial services firm at the heart of global metals trading.
- Benefits: Enjoy competitive pay, career growth, and a supportive work environment.
- Other info: Dynamic role with opportunities for professional development and leadership.
- Why this job: Make a real impact on critical infrastructure and drive innovation in reliability engineering.
- Qualifications: Experience in infrastructure reliability or site reliability engineering is essential.
The predicted salary is between 63000 - 77000 £ per year.
The Infrastructure Reliability Engineering Lead is responsible for leading and developing the Infrastructure Reliability Engineering team, ensuring the successful implementation of reliability engineering practices across the infrastructure platforms that underpin LME services. Working under the direction of the Infrastructure Reliability Engineering Senior Manager, the role translates Infrastructure Reliability Engineering strategy, standards and governance into practical engineering outcomes and measurable improvements in service reliability, resilience and operational risk reduction.
The role is accountable for driving engineering excellence across observability, resilience engineering, automation, operational readiness and service reliability. The postholder will lead a team of specialist Infrastructure Reliability Engineers, ensuring infrastructure services are designed, operated and continuously improved in a secure, recoverable, observable and supportable manner.
The Infrastructure Reliability Engineering Lead will work closely with Compute Service Operations, Information Security, Application Delivery and other technology teams to identify reliability risks, improve platform resilience, reduce operational toil through automation and ensure infrastructure services meet business, regulatory and operational requirements. The initial emphasis of the role will be to establish and mature Infrastructure Reliability Engineering practices, operational readiness assessments, resilience validation, platform observability standards and reliability-focused automation across critical infrastructure services.
Key responsibilities:
- Lead, coach and develop the Infrastructure Reliability Engineering team, setting clear direction, priorities, objectives and performance expectations.
- Build and maintain a high-performing engineering team through effective recruitment, performance management, coaching, mentoring and succession planning.
- Allocate engineering resources effectively across reliability improvement initiatives, resilience programmes, operational readiness activities and platform engineering priorities.
- Foster a culture of engineering excellence, accountability, continuous improvement and shared ownership of service reliability.
- Lead implementation and continual improvement of Infrastructure Reliability Engineering practices across infrastructure and platform services.
- Define, measure and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability metrics and reliability measures across critical infrastructure services.
- Identify, assess and drive remediation of reliability risks and systemic weaknesses before they impact production services.
- Establish and govern observability standards across infrastructure platforms including monitoring, telemetry, logging, tracing and alerting.
- Ensure meaningful operational visibility and actionable service health monitoring across infrastructure services.
- Lead improvements in monitoring quality, alert effectiveness and operational insights through data-driven analysis.
- Lead resilience validation activities including failover testing, disaster recovery testing, recovery assurance exercises and scenario-based resilience assessments.
- Establish and maintain Operational Readiness standards, ensuring services meet agreed supportability, recoverability, observability and operational acceptance criteria before entering production.
- Drive adoption of Infrastructure as Code, configuration management and automation practices across infrastructure services.
- Support the Infrastructure Reliability Engineering Senior Manager in implementing and maturing the Infrastructure Reliability Engineering operating model, standards and governance frameworks.
PERSON SPECIFICATION:
Academic and Professional Qualifications Required:
- Bachelor's degree in Computer Science, Engineering, Information Technology or a related technical discipline.
- Significant experience in Infrastructure Reliability Engineering, Site Reliability Engineering, Platform Engineering or Infrastructure Engineering disciplines.
- Demonstrable experience leading technical engineering teams within complex enterprise environments.
Required Knowledge and Level of Experience:
Candidates should have significant experience operating within complex enterprise technology environments, ideally within financial services or other highly regulated organisations, together with a strong track record of leading teams responsible for critical infrastructure services and engineering outcomes.
The successful candidate will demonstrate strong people leadership, sound technical judgement and the ability to influence engineering decisions across multiple technology domains. They will possess broad expertise in reliability engineering, resilience, observability, automation and platform engineering, while remaining capable of contributing hands-on where required.
Technical Skills set and Core Competencies Required for Role:
- Strong understanding of Infrastructure Reliability Engineering and Site Reliability Engineering principles and practices.
- Experience defining and improving SLIs, SLOs, availability targets and reliability metrics.
- Strong scripting/automation capability (e.g., Python, Bash, Ansible, Terraform, GitOps).
- Experience with CI/CD pipelines (GitHub Actions, Jenkins, Azure DevOps, GitLab, etc).
- Strong observability skills—including metrics, logs, distributed tracing, and tools such as Prometheus, Grafana, ELK, OpenTelemetry, Jaeger.
- Good understanding of enterprise infrastructure platforms including Linux, virtualisation, container platforms and cloud-native technologies.
- Strong understanding of resilience engineering, disaster recovery, and service continuity principles.
- Experience conducting root cause analysis and driving long-term reliability improvements.
- Understanding of operational readiness, service transition and supportability principles.
- Knowledge of infrastructure security, hardening, compliance controls and operational risk management.
- Understanding of ITIL-aligned Incident, Problem, Change and Service Management disciplines.
The following experience would be advantageous:
- Experience establishing or maturing Infrastructure Reliability Engineering, Site Reliability Engineering or Platform Engineering capabilities.
- Experience operating within financial services, exchanges, clearing organisations or other highly regulated environments.
- Experience supporting Kubernetes, OpenShift or cloud-native platform ecosystems.
- Experience implementing enterprise observability solutions and telemetry platforms.
- Experience leading operational readiness reviews and infrastructure service acceptance activities.
- Experience delivering large-scale infrastructure automation and standardisation initiatives.
- Experience working with third-party technology providers, managed service partners and infrastructure suppliers.
- Experience supporting significant infrastructure transformation or platform modernisation programmes.
Skills and competencies:
- Strong people leadership, coaching and team development skills.
- Strong analytical, troubleshooting and problem-solving capability.
- Excellent stakeholder management and relationship-building skills.
- Ability to provide technical leadership while maintaining focus on delivery, governance and business outcomes.
- Strong planning, prioritisation and organisational skills.
- Excellent written, verbal and presentation skills.
- Strong focus on automation, standardisation and continuous improvement.
- Data-driven and evidence-based approach to decision-making.
- Demonstrates ownership, accountability and engineering excellence.
- Comfortable operating within high-pressure, business-critical and regulated environments.
Infrastructure Reliability Engineering Lead in London employer: The London Metal Exchange Limited
The London Metal Exchange Limited is an exceptional employer, offering a dynamic work environment that fosters innovation and collaboration. With a strong emphasis on employee growth, the company provides ample opportunities for professional development and skill enhancement, particularly in data analytics and automation. Located in the heart of London, employees benefit from a vibrant city life while being part of a forward-thinking organisation committed to excellence in risk management and audit practices.
Contact Details:
The London Metal Exchange Limited Recruitment Team
StudySmarter Expert Advice🤫
We think this is how you could land Infrastructure Reliability Engineering Lead in London
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at The London Metal Exchange Limited or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to The London Metal Exchange Limited.
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like The London Metal Exchange Limited.
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like The London Metal Exchange Limited that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Infrastructure Reliability Engineering Lead in London
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at The London Metal Exchange Limited.
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at The London Metal Exchange Limited and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at The London Metal Exchange Limited
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If The London Metal Exchange Limited uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.