Remote AI Evaluation & Benchmark Engineer

Remote AI Evaluation & Benchmark Engineer

Full-Time Working from home possible
I

Intergral UK is seeking an AI Evaluation & Benchmarking Engineer to design and implement systems that measure the quality of OpsPilot. You will lead automated evaluations, create synthetic workloads, and benchmark models, prompts, tools, and workflows against repeatable baselines.

You will work with cross-functional teams to extend evaluation across customer journeys, APIs, and backend services, using OpenTelemetry to ensure representative benchmarks and correlate results with metrics, logs, and

#J-18808-Ljbffr

I

Contact Details:

Intergral GmbH Recruitment Team