Lead Site Reliability Engineer in London

Lead Site Reliability Engineer in London

London Full-Time 80000 - 100000 £ / year (est.) No working from home possible
J.P. Morgan

At a Glance

  • Tasks: Shape next-gen SRE patterns and enhance reliability across global trading systems.
  • Company: Join J.P. Morgan, a leader in financial services with a focus on innovation.
  • Benefits: Competitive salary, diverse culture, and opportunities for professional growth.
  • Other info: Be part of a dynamic team driving cutting-edge technology in finance.
  • Why this job: Make a real impact in a fast-paced environment while collaborating with traders.
  • Qualifications: Experience in high-pressure trading environments and strong programming skills.

The predicted salary is between 80000 - 100000 £ per year.

Our trading technology stack is undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization.

As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management's Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front-office trading platforms.

Job Responsibilities:
  • Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities.
  • Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs.
  • Support live trading environments, including incident response, root cause analysis, and post mortem leadership.
  • Work as a core member of the software engineering team, participating in daily standups and design discussions.
  • Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation.
  • Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self-healing workflows, and resilience engineering.
  • Use enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Lead reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Drive improvements in latency, throughput, and stability across high volume trading applications.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments.
  • Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC.
  • Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end-to-end reliability.
Required Qualifications, Capabilities, and Skills:
  • Strong hands-on experience in front office trading environments or similarly high pressure, low latency domains.
  • Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing and implementing observability frameworks for mission critical systems.
  • Proven ability to lead incident response and drive long term remediation.
  • Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code.
  • Experience with microservices, distributed systems, and event driven architectures.
  • Strong understanding of CI/CD pipelines, automated testing, and deployment strategies.
  • Comfortable interacting directly with traders and senior stakeholders.
  • Excellent communication skills, especially when translating technical issues into business impact.
  • Ability to operate calmly and decisively in high pressure situations.
  • Strong leadership presence with a collaborative mindset.

Lead Site Reliability Engineer in London employer: J.P. Morgan

As a leading global financial institution, we pride ourselves on fostering a dynamic and inclusive work environment that empowers our employees to excel. The Head of Markets Tax Operations role offers unparalleled opportunities for professional growth, collaboration across diverse teams, and the chance to drive innovation in tax operations while working in vibrant locations like Spain, Italy, and Hungary. Join us to be part of a culture that values excellence, embraces technology, and prioritises employee development, ensuring you can make a meaningful impact in your career.

J.P. Morgan

Contact Details:

J.P. Morgan Recruitment Team

We think you need these skills to ace Lead Site Reliability Engineer in London

Site Reliability Engineering (SRE)
Java
Kotlin
Python
FIX messaging
Kafka
Grafana