AI Evaluation & Benchmarking Engineer

AI Evaluation & Benchmarking Engineer

Full-Time 60000 - 75000 £ / year (est.) Working from home possible
Intergral

At a Glance

  • Tasks: Evaluate and benchmark AI systems to ensure quality and performance.
  • Company: Join a small, innovative tech company with a flat structure.
  • Benefits: Enjoy remote work, flexible hours, and 25 days holiday.
  • Other info: Collaborate closely with a small team and influence key decisions.
  • Why this job: Make a real impact on cutting-edge AI technology and drive continuous improvement.
  • Qualifications: 3+ years in AI or software engineering, strong Python skills required.

The predicted salary is between 60000 - 75000 £ per year.

Location: Remote (United Kingdom)

Salary: £60,000–£75,000 DOE

Contract: Full-time, 40 hours per week

Employer: Intergral UK

Reports to: Director of Engineering

You must already have the right to work in the UK, as we're unable to sponsor visas for this role.

About the Role

We're looking for an AI Evaluation & Benchmarking Engineer to determine how we measure the quality of OpsPilot, and to build the systems that do it. More important than any individual technology is how you approach measurement. We're looking for someone who questions whether a system is actually achieving the outcome it was designed for, works out how to measure that objectively, and builds what's needed to keep measuring it as the product changes. You'll have real autonomy over the technical approach. We set the goals and check in regularly, but the expertise on how to get there is yours — we're hiring you because we need someone who can take an ambiguous problem and deliver a working system without being handed the steps. This is a hands-on engineering role with no reports and no QA silo. Our engineering structure is flat: you'll report to the Director of Engineering and work alongside the other engineers as a peer. This isn't a traditional QA role. You won't be manually testing tickets, acting as a release gatekeeper or simply checking whether features technically work.

About OpsPilot

OpsPilot is an AI-led observability platform helping engineering and operations teams move from monitoring data to evidence-backed operational understanding — so they can investigate problems faster and act with greater confidence. Intergral has more than 20 years of experience in application performance and observability, with an established customer base built around FusionReactor. OpsPilot is expanding beyond its historic Java and ColdFusion roots into the wider observability market, supporting OpenTelemetry-based metrics, logs and traces. At the heart of OpsPilot is Coworker, an AI operations capability that continuously investigates telemetry, identifies situations that need attention and provides evidence-backed findings and recommended next steps.

What You'll Do

  • Evaluate the agent and the product it runs on. Coworker's findings are only as good as the system underneath them. An investigation can fail because the model reasoned badly, because retrieval surfaced the wrong evidence, because ingestion dropped a trace, or because the finding was presented in a way no engineer could act on. Evaluating the agent in isolation would tell us very little, so this role covers both.
  • On the agentic side, you'll:
    • Build automated evaluations for our agentic AI capabilities.
    • Create realistic synthetic scenarios, datasets and workloads with meaningful ground truth and evaluation criteria.
    • Measure task success, diagnostic accuracy, evidence quality, reliability, consistency, latency and cost.
    • Account for the non-deterministic nature of AI systems through repeated runs, variance analysis and determining whether changes are meaningful rather than noise.
    • Benchmark models, prompts, tools, retrieval strategies and agent workflows against repeatable baselines.
    • Use relevant industry benchmarks and standards, including established SRE practices and emerging AI-agent, AIOps and incident-response benchmarks, and build our own where existing approaches don't represent real operational work.
  • Across the wider product, you'll:
    • Extend evaluation across important customer journeys, APIs, backend services and UI.
    • Use OpenTelemetry, including its Semantic Conventions, to make benchmark environments representative and portable.
    • Use product telemetry and correlate benchmark results with metrics, logs and traces to understand why failures occur.
    • Build end-to-end measurements focused on customer outcomes rather than isolated components.
  • When we change a model, prompt, tool or agent workflow, we want to know what became better, what became worse and why — including the impact on quality, reliability, latency and cost.
  • Turn what we learn into continuous improvement.
  • Turn failures and real-world problems into new evaluation scenarios.
  • Identify recurring failure patterns and capability gaps.
  • Test potential improvements against established baselines, holdouts and unseen scenarios.
  • Detect regressions, benchmark overfitting and improvements that don't generalise.
  • Longer term, this evaluation system becomes the harness for controlled self-improvement: identifying weaknesses, testing potential changes and objectively determining whether they should be retained.
  • Work with the rest of engineering.
  • Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product.
  • Make evaluation failures clear, reproducible and actionable.
  • Make straightforward fixes yourself where that's the most efficient approach.
  • Build tooling that makes evaluations easy for other engineers to create, run and understand.
  • Use AI-assisted engineering where it improves the speed or quality of your work.
  • Benchmarking should provide continuous feedback that helps engineering improve the product. This role is not a release gatekeeper.

    What We're Looking For

    Above everything else: the ability to take an ambiguous technical problem, develop an approach and deliver a working system independently. We're more interested in demonstrated ability than an exact number of years, but we'd generally expect around 3+ years of relevant technical experience. Your background might be as an AI engineer, software engineer, SRE, platform engineer, performance engineer, SDET or similar. Alongside that, we're looking for:

    • Practical experience working with LLMs, AI agents or AI evaluation.
    • Strong software engineering skills, particularly Python or a similar language.
    • Experience building automated evaluation, benchmarking, testing or experimentation infrastructure.
    • Experience creating synthetic workloads, datasets or evaluation scenarios.
    • An understanding of non-deterministic evaluation, including repeated measurement, variance and distinguishing meaningful changes from noise.
    • The ability to turn complex system behaviour into measurable criteria.
    • Comfort working across APIs, distributed systems and multiple layers of a software product.

    Desirable

    Any of the following would be useful, but none are required:

    • Agentic AI evaluation, tool use and multi-step workflows.
    • LLM evaluation frameworks and model-based evaluation techniques.
    • Automated experimentation or self-improving systems.
    • Dataset, ground-truth and holdout evaluation design.
    • Statistical experimentation and performance benchmarking.
    • OpenTelemetry, metrics, logs and distributed tracing.
    • SRE, incident response or observability.
    • Production SaaS and distributed systems.

    What Success Looks Like

    We have the beginnings of an evaluation harness, but the design and expertise are what we're hiring for. We'd expect the first few months to go into the core evaluation harness and a starting corpus for Coworker's investigation quality, then extend outward across the rest of the product as that proves itself. How you sequence it is your call. By six months, we should be able to objectively answer questions such as:

    • Is Coworker getting better at investigating operational problems?
    • Where does it perform well or poorly, and why?
    • Are its conclusions supported by the right evidence?
    • How do different models, tools and agent configurations compare?
    • What quality, latency and cost trade-offs are we making?
    • How do we compare against relevant external benchmarks?
    • Have improvements introduced regressions elsewhere?
    • Where should we focus improvement next?

    Our evaluation corpus should keep growing as we encounter new problems. Success isn't measured by the number of tests written or percentage test coverage. It's measured by our ability to understand how well OpsPilot is doing its job, where it isn't, and whether the changes we're making are actually making it better.

    What We Offer

    A small company rather than a large one, with the trade-offs that implies. Under ten people in engineering, a flat structure, and decisions made in a conversation rather than across three meetings. You'll have genuine influence over how this is done, and very little bureaucracy to work through to get there.

    Fully remote within the UK. Flexible working hours. 25 days holiday plus bank holidays. Real autonomy over your technical approach and how you deliver the role.

    Our Interview Process

    Straightforward: usually two or three conversations, with no technical coding tests. If this sounds like the kind of challenge you're looking for, we'd love to hear from you.

    AI Evaluation & Benchmarking Engineer employer: Intergral

    OpsPilot is an exceptional employer, offering a fully remote work environment within the UK that promotes flexibility and genuine ownership in a small engineering team. With a modern cloud-native technology stack and minimal bureaucracy, employees are encouraged to innovate and influence product development directly, while benefiting from opportunities for professional growth and collaboration in a supportive culture.

    Intergral

    Contact Details:

    Intergral Recruitment Team

    StudySmarter Expert Advice🤫

    We think this is how you could land AI Evaluation & Benchmarking Engineer

    Join Local Tech Meetups

    Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Intergral or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

    Contribute to Open Source Projects

    Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Intergral.

    Tap into Online Developer Communities

    Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Intergral.

    Explore Job Boards Specifically for Tech Roles

    Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Intergral that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

    We think you need these skills to ace AI Evaluation & Benchmarking Engineer

    AI Evaluation
    Benchmarking
    Automated Evaluation
    Synthetic Workload Creation
    Data Analysis
    Python Programming
    Non-Deterministic Evaluation

    Some tips for your application 🫡

    Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

    Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Intergral.

    Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Intergral and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

    Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

    How to prepare for a job interview at Intergral

    Brush Up on Your Coding Skills

    For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

    Know Your Tools and Frameworks

    Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Intergral uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

    Showcase Your Projects

    Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

    Prepare for Behavioural Questions

    While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.