Member of Technical Staff – Environments / Evals
Stealth AI Lab | Paris or London (Hybrid)
About
AI models need useful problems to practise on if they're going to improve. This role is about creating those problems automatically — and then building evaluations that tell you whether the model has genuinely learned something.
The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.
You'll build systems that generate training tasks, adjust their difficulty as models improve, and simulate parts of real-world deployments. You'll also make sure improvements are real rather than models exploiting shortcuts in the training environment.
What you'll do
- Build pipelines that automatically generate tasks and training environments
- Create systems that filter tasks for validity, diversity and difficulty
- Adjust training curricula based on where models succeed and fail
- Train simulators using data from real model deployments
- Test how closely simulated environments match real interactions
- Build reproducible evaluation frameworks and held-out tests
- Identify reward hacking, shortcuts, contamination and failures to generalise
What you'll need
- Strong engineering skills and good experimental judgement
- Experience with PyTorch and modern ML systems
- Experience in synthetic data, agents, world models, curriculum learning or evaluation
- Ability to design rigorous model evaluations
- Strong understanding of model behaviour and failure analysis
- Comfort working across open-ended research and engineering problems
Optional
- Ray, Docker, OpenEnv or Gymnasium
- vLLM, SGLang, Playwright, Inspect AI or DeepEval
Shortlisted candidates will be contacted within 48 hours.
#J-18808-Ljbffr
Member of Technical Staff – RL Environments / Evals employer: Axiōma Search
Axiōma Search is an exceptional employer for those passionate about machine learning and time-series forecasting. With a collaborative work culture that prioritises innovation and employee growth, team members are encouraged to explore new ideas and technologies while benefiting from a global network of experts. Located in Europe, the company offers unique opportunities for professional development and the chance to make a significant impact in the field of AI.