Senior Research Scientist | Model Steering

Senior Research Scientist | Model Steering

Full-Time 70000 - 90000 £ / year (est.) No working from home possible
The Consensus

At a Glance

  • Tasks: Lead cutting-edge research in AI model steering and fine-tuning for translation models.
  • Company: Join DeepL, a global leader in AI technology with a mission to enhance communication.
  • Benefits: Enjoy flexible hours, hybrid work, competitive salary, and 30 days of annual leave.
  • Other info: Be part of a diverse team and enjoy regular team events and hack sessions.
  • Why this job: Shape the future of AI while working on impactful projects in a collaborative environment.
  • Qualifications: Expertise in LLM post-training, reinforcement learning, and strong coding skills required.

The predicted salary is between 70000 - 90000 £ per year.

Meet DeepL. DeepL is a global AI product and research company focused on building secure, intelligent solutions to complex business problems. Over 200,000 business customers and millions of individuals across 228 global markets today trust DeepL's Language AI platform for human-like translation, improved writing, and real-time voice translation.

Founded in 2017, DeepL now has around 1,000 passionate employees and is supported by world-renowned investors. Our goal is to become the global leader in trusted, intelligent AI technology, building products that drive better communication, foster connections, and create a meaningful impact. To achieve this, we need talented people like you to join our journey.

What sets us apart is our blend of cutting-edge AI technology, meaningful work, and a culture where people truly thrive. We’re a team of innovators, researchers, and creators driven by a shared purpose to unlock human potential by making work simpler, smarter, and more connected.

Our Language AI teams form the foundation of DeepL's success. We are a dedicated group of researchers who collaborate closely with engineers, product managers, and designers. Our mission is to build the world's leading language AI system to deliver perfect translations for the most demanding use cases. This includes data, training, quality assurance, and operational aspects.

Your responsibilities include:

  • Leading fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based translation models.
  • Driving hands-on research and development on post-training for our core translation models.
  • Building reward models and evaluator models for translation.
  • Owning the full lifecycle of model delivery.
  • Establishing strong practices for evaluation, reproducibility, monitoring, and continuous model improvement in production.
  • Mentoring researchers and engineers.

Qualities we look for:

  • Proven experience making large models steerable and instruction-following.
  • Deep, hands-on expertise in LLM post-training, knowledge distillation, and/or reinforcement learning.
  • Strong data-centric instincts for building synthetic-data and preference-data pipelines.
  • Experience designing evaluation and reward signals.
  • Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow).

Nice to have:

  • Demonstrated experience fine-tuning and training large models at scale.
  • Experience with machine translation, multilingual NLP, or language quality estimation.
  • Publications at top-tier venues.

What we offer:

  • Diverse and internationally distributed team.
  • Open communication, regular feedback.
  • Hybrid work, flexible hours.
  • Virtual Shares - An ownership mindset in every role.
  • Regular in-person team events.
  • Monthly full-day hacking sessions.
  • 30 days of annual leave.
  • Competitive benefits.

If this role and our mission resonate with you, but you’re hesitant because you don’t check all the boxes, don’t let that hold you back. At DeepL, it’s all about the value you bring and the growth we can foster together.

We are an equal opportunity employer. You are welcome at DeepL for who you are - we appreciate authenticity here. The more voices we have represented and amplified in our business, the more we will all succeed, contribute, and think forward!

Senior Research Scientist | Model Steering employer: The Consensus

At ClickHouse, we pride ourselves on being an excellent employer that fosters a collaborative and innovative work culture. Our remote-friendly environment in the United Kingdom allows for flexibility while providing ample opportunities for professional growth and development in the field of cloud infrastructure. Join us to be part of a team that values reliability, encourages continuous learning, and offers a supportive atmosphere where your contributions truly matter.

The Consensus

Contact Details:

The Consensus Recruitment Team

We think you need these skills to ace Senior Research Scientist | Model Steering

Model Fine-Tuning
Post-Training Techniques
Reinforcement Learning
Data Pipeline Development
Evaluation and Reward Signal Design
Large Language Model (LLM) Expertise
Python Programming