Lead Site Reliability Engineer

Lead Site Reliability Engineer

Full-Time 80000 - 100000 £ / year (est.) No working from home possible
PVH (Tommy Hilfiger/Calvin Klein)

At a Glance

  • Tasks: Lead the design of next-gen reliability frameworks and support live trading environments.
  • Company: Join JPMorgan Chase's innovative Trading Technology group.
  • Benefits: Competitive salary, diverse culture, and opportunities for professional growth.
  • Other info: Collaborate globally and drive improvements in high-volume trading applications.
  • Why this job: Make a real impact in a fast-paced trading environment with cutting-edge technology.
  • Qualifications: Experience in SRE, strong coding skills, and ability to thrive under pressure.

The predicted salary is between 80000 - 100000 £ per year.

Lead Site Reliability Engineer

Our trading technology stack is undergoing a multi-year convergence and modernization journey.

You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems.

This role is ideal for an SRE specialist who thrives in fast-paced front‑office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization.

As a Lead Site Reliability Engineer at JPMorgan Chase within J.

Morgan Asset Management's Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front‑office trading platforms.

  • Job Responsibilities
  • Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities.
  • Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs.
  • Support live trading environments, including incident response, root cause analysis, and post‑mortem leadership.

Work as a core member of the software engineering team, participating in daily standups and design discussions.

  • Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation.
  • Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‑healing workflows, and resilience engineering.
  • Use enterprise‑authorized AI capabilities within the work environment to accelerate major‑incident triage, troubleshooting, and post‑incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Lead reuse‑first adoption of AI‑assisted reliability workflows across SDLC/toolchain practices (e. g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Drive improvements in latency, throughput, and stability across high‑volume trading applications.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments.
  • Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC.
  • Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end‑to‑end reliability.
  • Required Qualifications, Capabilities, and Skills
  • Strong hands‑on experience in front‑office trading environments or similarly high‑pressure, low‑latency domains.
  • Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, Influx DB, MQ (IBM MQ or similar), Oracle DB.
  • Demonstrated experience using enterprise‑authorized AI capabilities within the work environment to improve SRE workflows (e. g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI‑assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Deep knowledge of reliability engineering principles: SLIs/SLOs, real‑time telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing and implementing observability frameworks for mission‑critical systems.
  • Proven ability to lead incident response and drive long‑term remediation.
  • Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production‑grade code.

Experience with microservices, distributed systems, and event‑driven architectures.

  • Strong understanding of CI/CD pipelines, automated testing, and deployment strategies.
  • Comfortable interacting directly with traders and senior stakeholders.

Excellent communication skills, especially when translating technical issues into business impact.

Ability to operate calmly and decisively in high‑pressure situations.

  • Strong leadership presence with a collaborative mindset.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success.

We are an equal opportunity employer and place a high value on diversity and inclusion at our company.

  • We do not discriminate on the basis of any prot
  • #J-18808-Ljbffr

Lead Site Reliability Engineer employer: PVH (Tommy Hilfiger/Calvin Klein)

Intapp is an exceptional employer, offering a dynamic work environment that fosters innovation and collaboration within the accounting and consulting sectors across EMEA. With a strong commitment to employee growth, Intapp provides ample opportunities for professional development and leadership coaching, ensuring that team members thrive in their careers while contributing to the company's strategic vision. The culture is built on accountability and high performance, making it an ideal place for those looking to make a significant impact in a rapidly evolving industry.

PVH (Tommy Hilfiger/Calvin Klein)

Contact Details:

PVH (Tommy Hilfiger/Calvin Klein) Recruitment Team

We think you need these skills to ace Lead Site Reliability Engineer

Site Reliability Engineering (SRE)
Java
Kotlin
Python
FIX messaging
Kafka
Grafana