Lead Site Reliability Engineer in London

Lead Site Reliability Engineer in London

London Full-Time 80000 - 100000 £ / year (est.) No working from home possible
Jpmorgan Chase & Co.

At a Glance

  • Tasks: Shape next-gen SRE patterns and enhance reliability across global trading systems.
  • Company: Join JPMorgan Chase's innovative Trading Technology group.
  • Benefits: Competitive salary, health benefits, and opportunities for professional growth.
  • Other info: Dynamic role with excellent career advancement opportunities and global collaboration.
  • Why this job: Make a real impact in a fast-paced trading environment while collaborating with traders.
  • Qualifications: Experience in high-pressure trading environments and strong programming skills in Python, Java, or Kotlin.

The predicted salary is between 80000 - 100000 £ per year.

Our trading technology stack is undergoing a multi‑year convergence and modernization journey.

You will play a pivotal role in shaping our next‑generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems.

This role is ideal for an SRE specialist who thrives in fast‑paced front‑office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization.

As a Lead Site Reliability Engineer at JPMorgan Chase within J.

Morgan Asset Management’s Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front‑office trading platforms.

Job Responsibilities

  • Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities.
  • Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs.
  • Support live trading environments, including incident response, root cause analysis, and post mortem leadership.

Work as a core member of the software engineering team, participating in daily standups and design discussions.

  • Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation.
  • Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‑healing workflows, and resilience engineering.
  • Use enterprise‑authorized AI capabilities within the work environment to accelerate major‑incident triage, troubleshooting, and post‑incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Lead reuse‑first adoption of AI‑assisted reliability workflows across SDLC/toolchain practices (e. g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Drive improvements in latency, throughput, and stability across high volume trading applications.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments.
  • Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC.
  • Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability.

Required Qualifications, Capabilities, and Skills

  • Strong hands‑on experience in front office trading environments or similarly high‑pressure, low‑latency domains.
  • Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, Influx DB, MQ (IBM MQ or similar), Oracle DB.
  • Demonstrated experience using enterprise‑authorized AI capabilities within the work environment to improve SRE workflows (e. g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI‑assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Deep knowledge of reliability engineering principles: SLIs/SLOs, real‑time telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing and implementing observability frameworks for mission critical systems.
  • Proven ability to lead incident response and drive long term remediation.
  • Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production‑grade code.

Experience with microservices, distributed systems, and event‑driven architectures.

  • Strong understanding of CI/CD pipelines, automated testing, and deployment strategies.
  • Comfortable interacting directly with traders and senior stakeholders.

Excellent communication skills, especially when translating technical issues into business impact.

Ability to operate calmly and decisively in high pressure situations.

  • Strong leadership presence with a collaborative mindset.
  • #J-18808-Ljbffr

Lead Site Reliability Engineer in London employer: Jpmorgan Chase & Co.

JPMorgan Chase & Co. is an exceptional employer, offering a dynamic work environment in Bournemouth where innovation and collaboration thrive. Employees benefit from a strong focus on professional development, inclusive team culture, and the opportunity to work with cutting-edge technology in a supportive atmosphere that values continuous improvement and engineering excellence.

Jpmorgan Chase & Co.

Contact Details:

Jpmorgan Chase & Co. Recruitment Team

We think you need these skills to ace Lead Site Reliability Engineer in London

Site Reliability Engineering (SRE)
Java
Kotlin
Python
FIX messaging
Kafka
Grafana