At a Glance
- Tasks: Lead the development of AI/ML infrastructure and optimise LLM inference performance.
- Company: Join JPMorganChase, a leader in financial technology innovation.
- Benefits: Attractive salary, health benefits, remote work options, and career growth opportunities.
- Other info: Dynamic team environment with cutting-edge technology and significant career advancement potential.
- Why this job: Make a real impact on AI capabilities at a global financial institution.
- Qualifications: Experience in software engineering, particularly with Python or Go, and LLM inference systems.
The predicted salary is between 72000 - 88000 £ per year.
At JPMorganChase, we are building the infrastructure that powers the next generation of enterprise AI — and we need talented engineers who are passionate about LLM inference to help us do it. This is your opportunity to work at the intersection of cutting-edge machine learning and large-scale production systems, directly contributing to how one of the world's largest financial institutions deploys and optimizes AI at scale.
As a Lead Software Engineer at JPMorganChase within the AI/ML Data Platform team, you will be a key technical contributor on LLM inference performance — supporting optimization strategy, benchmarking, and efficiency at scale. You will collaborate closely with senior engineers and engineering leadership to help shape how our platform evolves, ensuring every model we serve is fast, cost-efficient, and production-ready. This is a high-impact individual contributor role where your technical contributions will have direct, measurable influence on the firm's AI capabilities.
Job Responsibilities
- Execute systematic benchmarking and performance characterization across production LLM workloads, establishing reproducible baselines, identifying regressions, and quantifying the impact of configuration changes before they reach production.
- Design and run quantization experiments — FP8, INT8/INT4 (GPTQ/AWQ), and next-generation precision formats — measuring accuracy delta, throughput improvement, memory reduction, and cost-per-token impact.
- Support speculative decoding strategy across the model portfolio, including draft model, n-gram, and multi-token prediction approaches, contributing to acceptance rate measurement and per-workload configuration recommendations.
- Build and maintain GPU efficiency metrics covering utilization, memory headroom, cost per 1K tokens, and waste identification — providing engineering teams with a data-driven view of platform efficiency.
- Benchmark the platform against external providers and published industry numbers, identifying gaps and contributing to improvement initiatives.
- Participate in inference engine upgrade evaluations, including new scheduler architectures, async tensor parallelism, disaggregated prefill/decode, and advanced speculative decoding, supporting systematic validation before production promotion.
- Contribute to GPU chaos engineering efforts, including induced failure scenarios, hardware diagnostic monitoring, and detection and recovery measurement.
- Leverage enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards.
- Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and advanced applied experience – preferably Go / Python.
- Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving engines.
- Strong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus compute bottlenecks, and the practical implications of quantization at inference time.
- Experience with quantization techniques and their real-world tradeoffs at scale.
- Familiarity with speculative decoding and the variables that drive acceptance rates in production workloads.
- Rigorous benchmarking skills using GuideLLM, custom harnesses, or equivalent tooling, with the ability to support every performance claim with data.
- Experience operating in cloud GPU infrastructure at scale (AWS, Kubernetes-based managed inference services).
- Ability to communicate technical trade-offs clearly to engineering peers and senior stakeholders.
- Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, testing, troubleshooting, or documentation) with demonstrated ability to critically evaluate and validate AI-generated outputs.
- Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations.
Preferred qualifications, capabilities, and skills
- Experience with disaggregated prefill/decode serving architectures.
- Familiarity with GPU hardware diagnostics tools such as DCGM, NVML, or XID event tracking.
- Experience with ML observability and production monitoring for inference workloads.
- Awareness of the LLM inference competitive landscape with a track record of applying industry benchmarks to drive platform improvements.
Lead Software Engineer - Python / Go & AI/ML employer: JPMorgan Chase & Co.
JPMorgan Chase & Co. is an exceptional employer, offering a dynamic work environment in the heart of London’s International Private Bank. With a strong emphasis on professional development, employees benefit from comprehensive training programs and opportunities for career advancement, all while enjoying a collaborative culture that values teamwork and innovation. The role of Executive Assistant not only provides a chance to work closely with senior leaders but also allows for meaningful contributions to the success of the team in a prestigious financial institution.
StudySmarter Expert Advice🤫
We think this is how you could land Lead Software Engineer - Python / Go & AI/ML
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at JPMorgan Chase & Co. or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to JPMorgan Chase & Co..
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like JPMorgan Chase & Co..
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like JPMorgan Chase & Co. that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Lead Software Engineer - Python / Go & AI/ML
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at JPMorgan Chase & Co..
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at JPMorgan Chase & Co. and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at JPMorgan Chase & Co.
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If JPMorgan Chase & Co. uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.