Agentic Evaluation Engineer β€” Benchmarks & Systems

Agentic Evaluation Engineer β€” Benchmarks & Systems

Full-Time No working from home possible
A

Callosum in London is hiring for an expert to lead agentic evaluation. You’ll build principled benchmarks, evaluate heterogeneous models and hardware, and ensure results stand up to scrutiny across tasks and generations of silicon.

Your work informs product directions and system design in a fast-moving AI stack. You will design experiments, quantify uncertainty, and partner with engineering to deploy durable observability.

#J-18808-Ljbffr

Agentic Evaluation Engineer β€” Benchmarks & Systems employer: AI Startups UK

Lovable is an exceptional employer for data scientists, particularly in our dynamic London growth team. With a culture of extreme ownership and low-ego collaboration, we empower our employees to drive impactful growth metrics while working alongside talented engineers. Our commitment to innovation and rapid experimentation offers unparalleled opportunities for professional development, making Lovable a truly rewarding place to build your career.

A

Contact Details:

AI Startups UK Recruitment Team