Callosum is seeking a PhD-level researcher to own and build a unified benchmarking system for multi-step LLM work. You will design task suites, run sandboxed grading with real commits and traces, and base decisions on reproducible results rather than model scores.
Lead external benchmark co-publications, enforce contamination controls, and drive open, scalable evaluation. This London-based, in-person role offers equity, private healthcare, visa sponsorship, and relocation support.
#J-18808-Ljbffr