Callosum in London is hiring for an expert to lead agentic evaluation. Youβll build principled benchmarks, evaluate heterogeneous models and hardware, and ensure results stand up to scrutiny across tasks and generations of silicon.
Your work informs product directions and system design in a fast-moving AI stack. You will design experiments, quantify uncertainty, and partner with engineering to deploy durable observability.
#J-18808-Ljbffr
Agentic Evaluation Engineer β Benchmarks & Systems employer: AI Startups UK
Lovable is an exceptional employer for data scientists, particularly in our dynamic London growth team. With a culture of extreme ownership and low-ego collaboration, we empower our employees to drive impactful growth metrics while working alongside talented engineers. Our commitment to innovation and rapid experimentation offers unparalleled opportunities for professional development, making Lovable a truly rewarding place to build your career.