AI Benchmarking Engineer for Cross-Stack Evaluation

AI Benchmarking Engineer for Cross-Stack Evaluation

Full-Time No working from home possible
C

Callosum Technologies Ltd. is hiring for a research-driven role in London to build and run a unified benchmarking system for evaluating agentic and algorithmic LLM workflows.

You will design tasks, collect real commits/traces, and ensure robust, contamination-free evaluation with reproducible results. The role requires a PhD or equivalent track record, strong Python, and experience with sandboxed/distributed execution.

#J-18808-Ljbffr

AI Benchmarking Engineer for Cross-Stack Evaluation employer: Callosum Technologies Ltd

At Callosum Technologies Ltd., we pride ourselves on fostering a dynamic and innovative work culture in the heart of London. As a Senior API Platform Architect, you'll not only lead the architectural vision but also have ample opportunities for professional growth and development within a collaborative team environment. Our commitment to employee well-being is reflected in our competitive benefits package and the chance to work on cutting-edge technology that makes a real impact.

C

Contact Details:

Callosum Technologies Ltd Recruitment Team