Senior LLM Evaluation & Coding Engineer (Remote EU)

Senior LLM Evaluation & Coding Engineer (Remote EU)

Full-Time 63000 - 77000 Β£ / year (est.) Working from home possible
B

At a Glance

  • Tasks: Design coding tasks and evaluate AI model outputs to enhance reliability and quality.
  • Company: Join Braintrust, a leader in innovative software engineering and AI evaluation.
  • Benefits: Flexible remote work, competitive pay, and opportunities for professional growth.
  • Other info: Collaborative team environment with exciting challenges and career advancement.
  • Why this job: Make a real impact on AI development while working with cutting-edge technology.
  • Qualifications: Experience in software engineering and a passion for AI and model evaluation.

The predicted salary is between 63000 - 77000 Β£ per year.

Braintrust is hiring experienced software engineers to join our evaluation and annotation team for a contracting engagement.

The role focuses on real-world software engineering, model evaluation, and applied AI, aiming to improve model reliability, reasoning, and code quality.

You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.

#J-18808-Ljbffr

Senior LLM Evaluation & Coding Engineer (Remote EU) employer: Braintrust

Braintrust is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the rapidly evolving AI landscape. As a Remote Regional Sales Director in EMEA, you will benefit from flexible time off, competitive salary and equity, and a supportive environment that prioritises employee growth and development. Join us to lead a high-performing sales team and make a significant impact in shaping the future of AI observability while enjoying perks like daily lunches and an AI stipend.

B

Contact Details:

Braintrust Recruitment Team

We think you need these skills to ace Senior LLM Evaluation & Coding Engineer (Remote EU)

Software Engineering
Model Evaluation
Applied AI
Coding Task Design
Benchmarking
Failure Mode Identification
Reinforcement Learning