About the Client
We’re working with an early-stage, well-funded AI startup (backed by top-tier VCs) building systems that need to understand and control how complex, large-scale AI behaviour plays out in the real world - before it goes wrong. The work sits right at the intersection of ML performance and safety: models need to be capable, but also predictable, aligned, and robust once they’re live in front of real customers.
Small team, high ownership, direct access to founders. This is a founding/early hire, not a cog-in-a-machine role.
The Role
The core of this role is RL, fine-tuning, and reward design -with safety as the design constraint, not an afterthought. You’ll take training methods and turn them into systems that are fast and effective, but also well-behaved: models that stay within intended bounds, resist drift, and fail safely rather than silently.
You’ll own the loop end-to-end: reward/training design, post-training and distillation, and production optimisation - all with an eye on catching and correcting unwanted behaviour before it reaches a customer.
What You’ll Do
- Design reward functions and training setups that optimise for capability and safe, predictable behaviour
- Post-train, fine-tune, and distil models with alignment and robustness front of mind
- Build evaluation and monitoring approaches that catch drift, edge cases, and failure modes early
- Optimise inference for scale without compromising on safety guardrails - sub-second, high-volume, production-grade
- Build the data pipelines that feed training, evaluation, and safety testing
- Take a method from prototype to production, simplifying aggressively while preserving the safety properties that matter
What We’re Looking For
- 2+ years shipping ML in a startup environment, ideally with end-to-end ownership
- Strong hands-on experience with RL, fine-tuning, and/or distillation - bonus points if you’ve thought hard about reward hacking, alignment, or failure modes
- Comfortable with real-time/low-latency inference systems
- Fluent with statistics, probability, and high-dimensional reasoning
- A genuine interest in safety-conscious ML, not just raw performance chasing
- Fast, high-bar, ownership mentality
Nice to have: distributed training experience, distilling frontier models into small open-weights models for production, background in anomaly/behavioural detection, interpretability or evals work, familiarity with cloud infra (AWS/GCP).
#J-18808-Ljbffr
Founding Engineer - AI Safety & Optimisation employer: Few&Far
Join a pioneering AI and robotics startup in London as a Founding Engineer, where you'll have the unique opportunity to shape cutting-edge technology from the ground up. With a strong emphasis on collaboration and innovation, this role offers genuine ownership and the chance to work alongside an exceptional founding team, backed by world-class investors. The vibrant work culture fosters personal growth and encourages tackling complex engineering challenges, making it an ideal environment for those looking to make a meaningful impact in the tech industry.