At a Glance
- Tasks: Optimise AI workloads and enhance performance across cutting-edge systems.
- Company: Join Lightning AI, a leader in AI infrastructure and innovation.
- Benefits: Enjoy competitive salary, equity, unlimited PTO, and wellness perks.
- Other info: Flexible work options and a culture of continuous improvement.
- Why this job: Make a real impact in AI while collaborating with top talent.
- Qualifications: Strong skills in deep learning frameworks and software engineering.
The predicted salary is between 63000 - 77000 £ per year.
Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.
We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London.
The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:
- Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum.
- Take Ownership: We own outcomes, not just our individual work.
- Communicate Openly: We communicate directly, seek to understand, and create clarity for others.
- Build Great Teams: We lead by example, empower others, and create healthy teams.
- Raise the Bar: We're always improving ourselves.
- Think Long-Term: We design for what's next.
We are seeking a highly skilled Research Engineer to help optimize training and inference workloads running on Lightning AI infrastructure. This role sits at the intersection of ML systems, AI infrastructure, performance engineering, and practical research.
This is a highly cross-functional role that combines deep technical problem solving with hands-on implementation. Successful candidates are comfortable working broadly across the stack while collaborating closely with customers and internal engineering teams to solve complex AI performance challenges.
This role can be based in one of our hubs (NYC, SF, Seattle, or London) or remote, with a minimum of 2 in-office days per week and occasional team and company offsites.
What You'll Do:
- Optimize large-scale training and inference workloads across GPUs, accelerators, and distributed systems.
- Work directly with customers to analyze workloads, identify bottlenecks, and improve performance, scalability, and reliability of deployed AI systems.
- Develop and improve inference pipelines, model serving systems, and performance-oriented tooling for production AI workloads.
- Design and implement profiling, debugging, and observability tools to analyze model execution and guide optimization strategies.
- Work across the software stack to ensure performance improvements are accessible through clean APIs, automation, and seamless integration with the Lightning ecosystem.
- Partner with hardware vendors and ecosystem partners to support efficient execution across diverse compute backends.
- Contribute to open-source projects through new features, tooling improvements, documentation, and community engagement.
- Stay current with advancements in large-scale inference, distributed training, and ML systems optimization.
What You’ll Need:
- Strong expertise with deep learning frameworks such as PyTorch.
- Experience working with large-scale training or inference workloads.
- Familiarity with distributed systems and parallelism strategies.
- Strong software engineering fundamentals.
- Experience analyzing and improving performance bottlenecks in ML systems.
- Excellent collaboration and communication skills.
- Ability to work comfortably in ambiguous, fast-moving environments.
- Bachelor’s degree in Computer Science, Engineering, or a related field.
Nice-to-Haves:
- Experience with inference optimization techniques.
- Experience with technologies such as CUDA, Triton, TensorRT, vLLM, SGLang, Dynamo, or related ML systems/inference tooling.
- Experience contributing to open-source ML, infrastructure, or AI systems projects.
- Startup experience or experience working in highly cross-functional environments.
- Advanced degree (Master’s or PhD) in AI, machine learning, systems, or related fields.
We are committed to offering competitive compensation that reflects the value each team member brings to our mission. Final offers are based on factors such as experience, skills, geographic location, and role expectations.
Benefits and Perks:
- Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
- Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
- Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
- Flexible Time Off: Unlimited PTO, company holidays, and floating holidays.
- Company-Wide Winter Break: Two weeks of company closure each winter.
- Paid Parental & Family Leave: Paid leave to support you and your family.
- Professional Development: Annual learning and development allowance.
- Wellness Benefits: Wellness and work-from-home stipends.
- Sabbatical Program: Four weeks of paid sabbatical leave after four years of service.
- Flexible Work: Flexible schedules and a hybrid work model.
- In-Office Meals: Complimentary meals at our office hubs.
At Lightning AI, we are committed to fostering an inclusive and diverse workplace. We provide equal employment opportunities to all employees and applicants.
Research Engineer employer: Lightningai
Lightning AI is an exceptional employer that champions innovation and collaboration in the heart of major tech hubs like London. With a strong focus on employee growth, we offer comprehensive benefits including unlimited PTO, meaningful equity, and professional development allowances, all within a dynamic work culture that values ownership and open communication. Join us to be part of a diverse team that not only drives cutting-edge AI solutions but also prioritises your well-being and career advancement.