AI Benchmark Scientist β€” Computational Mathematics

AI Benchmark Scientist β€” Computational Mathematics

Full-Time No working from home possible
M

Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks (Sci Code) for a 6-week, part-time engagement in London. You will design original, executable research problems for frontier models and calibrate scoring against model failures.

Responsibilities include sourcing materials, writing prompts, building grading criteria, and testing runs in a pull-request workflow with Docker/Git. Start date is immediate; ~20+ hours per week with flexible scheduling and remote-friendly setup

#J-18808-Ljbffr

AI Benchmark Scientist β€” Computational Mathematics employer: Mercor

Mercor is an exceptional employer that champions innovation and creativity in the AI sector, offering a fully remote work environment that promotes flexibility and independence. With a strong focus on employee growth, team members are encouraged to take ownership of their projects while benefiting from a supportive culture that values collaboration and continuous learning. Joining Mercor means being part of a forward-thinking company backed by industry leaders, where your contributions directly impact cutting-edge search technologies.

M

Contact Details:

Mercor Recruitment Team