Contract AI Evaluation Engineer β€” Coding Agent Tests in England

Contract AI Evaluation Engineer β€” Coding Agent Tests in England

England Full-Time 50000 - 70000 Β£ / year (est.) No working from home possible
Mindrift

At a Glance

  • Tasks: Create and evaluate AI coding tests in realistic simulated environments.
  • Company: Mindrift connects specialists with top tech companies for exciting AI projects.
  • Benefits: Flexible work hours, competitive pay, and opportunities for skill development.
  • Other info: Dynamic role with potential for growth in the tech industry.
  • Why this job: Join a cutting-edge team and shape the future of AI technology.
  • Qualifications: Experience in AI evaluation and strong problem-solving skills.

The predicted salary is between 50000 - 70000 Β£ per year.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.

You will create challenging tasks and evaluation criteria within realistic simulated environments.

We are building a dataset to evaluate AI coding agents and design tests that accept all valid approaches while rejecting incorrect ones.

This role involves guiding and evaluating agent solutions and iterating based on QA feedback.

#J-18808-Ljbffr

Contract AI Evaluation Engineer β€” Coding Agent Tests in England employer: Mindrift

Mindrift is an excellent employer for those seeking a flexible and rewarding role in the innovative field of AI mortgage underwriting. With competitive compensation and a focus on part-time project-based work, employees benefit from a supportive work culture that values expertise and encourages professional growth. Located in Edinburgh, Mindrift offers a unique opportunity to contribute to cutting-edge technology while maintaining a healthy work-life balance.

Mindrift

Contact Details:

Mindrift Recruitment Team

We think you need these skills to ace Contract AI Evaluation Engineer β€” Coding Agent Tests in England

AI Evaluation
Task Creation
Evaluation Criteria Development
Simulated Environment Design
Dataset Building
Coding Agent Testing
Quality Assurance Feedback Integration