Research Associate in Adaptive and Efficient LLM Architectures in London

Research Associate in Adaptive and Efficient LLM Architectures in London

London Full-Time 45399 - 48876 £ / year (est.) No working from home possible
I

At a Glance

  • Tasks: Lead transformative research in adaptive AI architectures and publish groundbreaking findings.
  • Company: Join the prestigious Imperial College London and work with top researchers.
  • Benefits: Competitive salary, extensive conference funding, and access to cutting-edge technology.
  • Other info: Flexible start date, diverse work culture, and excellent career growth opportunities.
  • Why this job: Make a real impact in the exciting field of generative AI and efficient architectures.
  • Qualifications: PhD in computer science or equivalent experience, strong coding skills, and a passion for research.

The predicted salary is between 45399 - 48876 £ per year.

We are looking for creative and passionate researchers to join the ERC project AToM (Adaptive Tokenization and Memory in Foundation Models) at Imperial College London, led by Dr. Edoardo Ponti, in a fully funded postdoctoral role to lead transformative research in adaptive and efficient architectures for AI models.

The Department of Computing is seeking highly motivated and talented Postdoctoral Research Associates, who have conducted cutting-edge research and/or have extensive experience in frontier AI labs. The position is fully funded by the ERC project with a focus on designing, implementing, and publicly releasing LLMs with efficient architectures (including adaptive memory, latent tokenization, sparse attention, multi-token prediction, adaptive depth, among others).

The recent revolution in generative AI has been driven by the rapid growth in training and inference compute for Foundation Models (FMs). This scaling paradigm, however, is characterised by high energy demand, high latency, and environmental impact. AToM sets out to reverse this trend by targeting a fundamental inefficiency in dominant FM architectures such as Transformers (including SWA/SSM hybrids).

Currently, models store, access, and convert data into sequences of internal representations, whose length bottlenecks both training/prefill (compute-bound) and decode (memory-bandwidth bound). Yet the granularity of the representations is largely determined upfront by input segmentation (tokenisation), which typically remains fixed across layers, and by the memory update mechanism, which accumulates most tokens in the key–value cache.

AToM aims to prototype new classes of FM architectures that learn, end-to-end, to compress their sequences of internal representations, effectively redefining the model’s “atomic units” for processing and memorising information. To accelerate adoption, we will retrofit state-of-the-art open-weight FMs into adaptive variants. In the first year of the project, we have already successfully released Qwen 3 with adaptive memory in collaboration with NVIDIA, and OLMo 3 with latent tokenization in collaboration with AI2.

This project will lead not only to substantial gains in efficiency (several orders of magnitude speedups without accuracy degradation) but also to the emergence of new capabilities: adaptive FMs can operate over broader effective horizons, as they can perceive longer inputs and generate longer outputs under a fixed budget. This enables (1) lifelong learning via a permanent, sub-linearly growing memory, (2) inference-time hyper-scaling for reasoning and agentic tasks (e.g., advanced maths, science, coding), and (3) world modelling for multimodal planning and simulation. Adaptive FMs thus open a path towards more efficient and more capable generative AI.

Within the project, you will conduct original research in the new and exciting field of efficient and adaptive LLM architectures and explore its applications across long-context understanding and reasoning (for code, maths, and agentic workflows) as well as long-horizon multimodal world modelling. We will strive to release new, more efficient and capable AI models and to publish in top-tier conferences and journals. Specifically, you will work closely with Dr. Edoardo Ponti in:

  • Designing and developing novel architectures for AI models
  • Performing retrofitting, post-training, and evaluation of SOTA open-weight models
  • Conducting independent and collaborative research within the group
  • Publishing results in top-tier conferences and journals (e.g., NeurIPS, ICML, ICLR, *ACL, EMNLP, Nature)
  • Implementing research ideas using modern deep learning frameworks (PyTorch/JAX), model/dataset libraries (Huggingface transformers/diffusers), and efficient kernels (Triton/CUDA)
  • Contributing to research projects on adaptive memory, latent tokenization, sparse attention, long-context understanding and reasoning, agentic, and multimodal world modelling
  • Contributing to the life and development of Edoardo Ponti’s Lab, including weekly meetings, presentations, maintaining the website and other resources
  • Delivering tutorials at conferences and summer schools
  • Mentoring PhD and MSc students
  • Presenting research at international conferences and workshops
  • Contributing to grant proposals and collaborative research initiatives

We are looking for self-driven and motivated individuals with genuine love for research. The applicant is also expected to have a strong track record in top conferences and journals in the fields of AI/ML/NLP, such as NeurIPS, ICML, ICLR, *ACL, EMNLP, etc. Candidates with research experience as part of frontier AI labs are also welcome.

Excellent skills in coding, strong foundations in mathematics (especially information theory, linear algebra, calculus), and knowledge in deep learning are essential. Practical experience in a broad range of techniques including LLM training, evaluation, RLVR, PEFT, quantisation, tensor/data parallelism is required. Ideally, the candidate should have familiarity with CUDA kernels and/or Triton, and inference engines (vLLM, SGLang, et cetera). Experience coding with deep learning libraries such as Pytorch/JAX is essential. Fluent written and spoken English skills as well as contributions to the group culture are expected. Applicants must hold a PhD in computer science or equivalent experience.

Extensive funding for conference travel (2 international conferences per year) and access to extensive compute via the AToM project GPUs (B200s and cloud credits) and GPUs from the Department of Computing and the College of Engineering (A100s and H200s) are offered. You will have the opportunity to continue your career at a world-leading institution and be part of our mission to continue science for humanity. Grow your career: gain access to Imperial’s sector-leading dedicated career support for researchers as well as opportunities for promotion and progression. A sector-leading salary and remuneration package (including 43 days off a year and generous pension schemes) is provided. Be part of a diverse, inclusive and collaborative work culture with various staff networks and resources to support your personal and professional wellbeing.

Full-time, fixed term 2-year contract to start around winter 26/27 (the start date is flexible). Candidates who have not yet been officially awarded their PhD will be appointed as Research Assistant within the salary range 45,399 - 48,876 per annum.

In addition to completing the online application candidates should attach a full CV with a list of all publications and a research statement (max 2 pages) indicating what you see are the most interesting research questions relating to adaptive and efficient AI architectures and why your expertise is relevant.

Informal enquiries related to the position should be directed to Dr Edoardo Ponti: eponti@ic.ac.uk. For queries regarding the application process contact Jamie Perrins: j.perrins@imperial.ac.uk.

Research Associate in Adaptive and Efficient LLM Architectures in London employer: Imperial College London

Imperial College London is an exceptional employer, offering a vibrant work culture that fosters collaboration and innovation within the International Relations Office. Employees benefit from professional growth opportunities, a commitment to diversity, and the chance to engage in meaningful projects that have a global impact, all set in the heart of one of the world's leading academic institutions.

I

Contact Details:

Imperial College London Recruitment Team

We think you need these skills to ace Research Associate in Adaptive and Efficient LLM Architectures in London

Research Skills
Deep Learning Frameworks (PyTorch/JAX)
Mathematics (Information Theory, Linear Algebra, Calculus)
LLM Training and Evaluation
Adaptive Memory Techniques
Latent Tokenization
Sparse Attention Mechanisms