Lead Software Platform Engineer in Glasgow

Lead Software Platform Engineer in Glasgow

Glasgow Full-Time No working from home possible
H

Help shape how AI systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. As a Lead Software Engineer at JPMorgan Chase in the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability, and lead incident response and continuous improvement.
\Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure

Build backend services and APIs that enable reliable operation of AI infrastructure in production

Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines

Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization techniques (e.g., Lead reliability engineering for LLM endpoints through capacity planning, load/soak testing, safe rollouts (blue/green, canary), failover, and incident response for outages and model-quality regressions

Participate in an on-call rotation, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-ups

Identify recurring operational issues and automate remediation to improve platform stability and developer experience

Build and maintain multi-agent systems with strong orchestration (planning, coordination, tool-calling, state/memory, and workflow control) where appropriate

Contribute to an inclusive team culture grounded in diversity, opportunity, inclusion, and respect, and help drive adoption of leading-edge technologies through communities of practice
\Formal training, certification, or equivalent practical experience in software engineering concepts

Hands-on experience with system design, application development, testing, and operational stability in production environments

Advanced proficiency in Python for building production-grade services and tooling

Hands-on experience with AWS and Terraform for infrastructure delivery and lifecycle management

Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns

Hands-on production experience operating LLM inference servers such as vLLM and llm-d (or directly equivalent serving stacks)

Hands-on experience hosting and serving LLMs on Amazon EKS and/or Amazon SageMaker, and on local GPU infrastructure

Knowledge of LLM reliability and risk considerations, including latency/throughput trade-offs, model and weight versioning, prompt/response logging, and safe rollout patterns
\Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns

Experience building AI agents using frameworks such as LangChain, CrewAI, LangGraph, or similar orchestration platforms

Experience operating or integrating serving platforms such as KServe, Ray Serve, NVIDIA Triton Inference Server, Text Generation Inference (TGI), alongside vLLM/llm-d

Experience with online LLM quality monitoring (e.g., Contributions to open-source LLM serving or inference projects (e.g., Morgan is a global leader in financial services, providing strategic advice and products to the world's most prominent corporations, governments, wealthy individuals and institutional investors. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives.
\We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. for more information about requesting an accommodation.
\Our professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing.

Lead Software Platform Engineer in Glasgow employer: Hackajob Ltd

JPMorgan Chase is an exceptional employer, offering a dynamic work environment where innovation and collaboration thrive. As a Lead Site Reliability Engineer, you will not only tackle complex challenges but also benefit from extensive professional development opportunities and a strong commitment to diversity and inclusion. Located in a global financial hub, you'll be part of a team that values your expertise and encourages a culture of continuous improvement and technical excellence.

H

Contact Details:

Hackajob Ltd Recruitment Team