Overview
In this leadership role, you steer the post-training research program for Thomson's LLMs, shaping data strategies, training, and evaluation to deliver accurate, well-calibrated models for regulated use. You collaborate with experimental, infrastructure, and evaluation teams to turn research into production-ready AI that supports professional decision-making. The role offers influence over roadmap priorities and exposure to leading AI methods and open-source contributions. You will publish findings and drive impactful, real-world AI applications in high-stakes domains.
Pay / Benefits
- hybrid work model
- flexible vacation
- Headspace app access
- tuition reimbursement
- volunteer days
- wellbeing resources
Responsibilities
- Oversee hands-on post-training for Thomson's LLMs: supervised fine-tuning, preference optimization (e.g., DPO), and reinforcement learning including agentic, multi-step settings with tool use.
- Establish and operate online, agentic reinforcement-learning pipelines with subject-matter experts in the loop.
- Own data selection, mixture optimization, synthetic-data generation, and evaluation design with measurable effects linked to model behavior.
- Detect training issues early and determine corrective actions.
- Collaborate daily with infrastructure and evaluation teams to maintain reliable training and evaluation pipelines.
- Provide actionable findings to inform roadmap and prioritization decisions.
Key requirements
- Post-training and reinforcement learning depth with hands-on experience in supervised fine-tuning, preference optimization, and RL in agentic, multi-step contexts.
- Data-centric model development with demonstrated data selection, mixture design, synthetic-data generation, and evaluation design.
- Engineering depth in distributed training, data pipelines, and evaluation infrastructure; ability to build and debug systems.
- Track record via shipped models, open-source contributions, or peer-reviewed publications at top venues (e.g., NeurIPS, ICML, ACL, EMNLP).
- Technical leadership experience heading a focused training or evaluation team.
- technical leadership
- cross-functional collaboration
- strong communication
- distributed training
- data pipelines
- evaluation infrastructure
Director, Model Research & Development in London employer: Thomson Reuters
Thomson Reuters is an exceptional employer, offering a dynamic work environment where innovation thrives and employees are empowered to lead the development of cutting-edge AI solutions. With a strong commitment to work-life balance, comprehensive benefits, and a culture that prioritises inclusion and professional growth, employees can expect to make a meaningful impact while advancing their careers in a globally recognised organisation. The hybrid work model and opportunities for community engagement further enhance the appeal of working in London, a vibrant hub for technology and legal expertise.