Callosum in London is building a unified benchmarking system for evaluating agentic and algorithmic AI solutions. You will design a harness to measure task success, quality, and robustness across motifs, agent topologies, and decomposition strategies, grounded in real commits and sandboxed grading.
This research hire builds and curates the evaluation framework, delivering proof points for customers and contributing to co-published benchmarks.
#J-18808-Ljbffr