At a Glance
- Tasks: Own the performance of our inference stack and optimise model serving efficiency.
- Company: Join a forward-thinking AI company focused on real-time adaptive intelligence.
- Benefits: Flexible work, travel stipend, lunch allowance, and comprehensive medical benefits.
- Other info: Collaborative environment with opportunities for personal and professional growth.
- Why this job: Make a real impact in AI by optimising performance and shaping the future of intelligence.
- Qualifications: 5+ years in ML systems, strong Python skills, and experience with performance engineering.
The predicted salary is between 72000 - 88000 £ per year.
The role
You'll own the cost and performance of our inference stack. Your work will determine how efficiently we serve models as workloads, traffic, and hardware change.
You'll work closely with the engineers operating the serving fleet while owning the core performance levers: caching, batching, quantization, decoding, and kernel-level optimization.
Success means improving throughput and latency without compromising reliability or model quality.
Responsibilities
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads based on real production traffic.
- Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
- Work within serving engines such as v LLM, SGLang, and Tensor RT-LLM, going below the framework when needed.
- Build profiling and measurement systems that show where time, memory, and compute are being spent.
Qualifications
- 5+ years in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with serving engines such as v LLM, SGLang, or Tensor RT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
Above all, we're looking for great teammates who make work feel lighter and aren't afraid to go out on a limb with bold ideas.
You don't need to be perfect, but you do need to be adaptable.
About Us
Most AI is frozen in place - it doesn't adapt to the world.
We think that's backwards.
Our mandate is to build efficient intelligence that evolves in real-time.
Our vision is AI systems that are flexible, personalized, and accessible to everyone.
We believe efficiency is what makes this possible - it's how we expand access and ensure innovation benefits the many, not the few.
We believe in talent density: bringing together the best and most driven individuals to push the boundaries of continual adaptation.
We're looking for builders and creative thinkers ready to shape the next era of intelligence.
Benefits
- Flexible work: In-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
- Adaption
Passport: Annual travel stipend to explore a country you've never visited.
We're building intelligence that evolves alongside you, so we encourage you to keep expanding your horizons.
- Lunch Stipend: Weekly meal allowance for take-out or grocery delivery.
- Well-Being: Comprehensive medical benefits and generous paid time off.
- #J-18808-Ljbffr
Inference Performance Engineer employer: Adaption
As a Senior Research Scientist at our innovative AI company, you'll be part of a dynamic team that values technical excellence and real-world impact. We offer a flexible work environment in the Bay Area, comprehensive medical benefits, and unique perks like an annual travel stipend to broaden your horizons. Join us to collaborate with top talent and contribute to cutting-edge research that shapes the future of intelligent systems.
StudySmarter Expert Advice🤫
We think this is how you could land Inference Performance Engineer
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Adaption!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Inference Performance Engineer at Adaption.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Adaption.
✨Apply Directly through Our Website
When you find a suitable opening like Inference Performance Engineer at Adaption, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Inference Performance Engineer
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at Adaption, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Adaption. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at Adaption
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Adaption!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.