Lead Site Reliability Engineer

Lead Site Reliability Engineer

Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
Hackajob Ltd

At a Glance

  • Tasks: Lead initiatives to enhance application reliability and mentor fellow engineers.
  • Company: Join JPMorgan Chase, a globally recognised firm with a focus on innovation.
  • Benefits: Competitive salary, professional development, and opportunities for career advancement.
  • Other info: Dynamic team environment with a strong emphasis on collaboration and growth.
  • Why this job: Make a significant impact in site reliability while working with cutting-edge technology.
  • Qualifications: Experience in site reliability engineering and proficiency in Python required.

The predicted salary is between 63000 - 77000 Β£ per year.

hackajob is partnering directly with JPMorganChase to hire for this role.

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them.

Job responsibilities:

  • Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team.
  • Leads initiatives to improve the reliability and stability of your team's applications and platforms using data-driven analytics to improve service levels.
  • Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers.
  • Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology-related bottlenecks in your areas of expertise.
  • Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses.
  • Documents and shares knowledge within your organization via internal forums and communities of practice.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills:

  • Formal training or certification on site reliability engineering concepts and advanced applied experience.
  • Should be able to design and code complex problems in Public cloud like AWS.
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform.
  • Fluency in Python.

Lead Site Reliability Engineer employer: Hackajob Ltd

At loveholidays, we pride ourselves on fostering a collaborative and innovative work culture that empowers our employees to thrive. As a Product Designer, you'll have the opportunity to contribute to meaningful projects that enhance customer experiences while enjoying a range of benefits, including professional development opportunities and a supportive team environment in a vibrant location. Join us in our mission to make travel accessible for everyone and be part of a company that values your creativity and input.

Hackajob Ltd

Contact Details:

Hackajob Ltd Recruitment Team

We think you need these skills to ace Lead Site Reliability Engineer

Site Reliability Engineering
Leadership Skills
Technical Expertise
Data-Driven Analytics
Service Level Indicators
Incident Management
Problem-Solving Skills