Machine Learning Engineer (LLM & Agent Evaluation) Mid to senior ML Engineering role at a fast-scaling, AI-first software business Build scalable evaluation infrastructure and multi-agent architectures at the core of the product Belfast based, hybrid with async-friendly global team Salary: competitive, reflecting experience, with equity UK work authorisation required About the Company Our client is a fast-scaling AI software business powering enterprise automation for Fortune 500 clients including major names across financial services and healthcare technology. Their small, elite Data Science and AI team builds and deploys cutting-edge ML and agentic AI systems at scale, with a culture built around intellectual curiosity, hands-on leadership and pragmatic startup thinking. Leaders stay close to the code, debate ideas openly and move fast without corporate inertia. This is a team where exceptional engineers thrive. The Role A newly created individual contributor position for an ML Engineer who wants to own evaluation end to end. You will design robust evaluation frameworks, build automated scoring and regression testing pipelines, and track quality across model, prompt and agent behaviour changes over time. A core part of this role involves building the infrastructure that converts expensive frontier agent tokens into optimised internal neural inference, a genuine and proprietary competitive advantage. Working closely with engineering teams and data scientists, you will analyse edge-case failure modes, build real-time quality dashboards and ensure high-confidence deployment workflows as the system scales. Key Responsibilities Design evaluation frameworks and metrics covering accuracy, safety, latency and cost across agent and LLM systems Build automated scoring pipelines, rubric-based grading and LLM-as-judge systems that scale beyond manual review Design and build chained multi-agent architectures that underpin the core product capability Stand up automated regression suites that catch quality drops from model, prompt or agent-logic changes before they reach production Bridge high-cost frontier models into low-cost, optimised internal inference systems Build dashboards and reporting that track model and agent quality over time and across releases Identify failure modes and edge cases across diverse scenarios, working with the team to prioritise fixes Partner closely with engineers building agent capabilities and with the senior data scientist for deeper analytical support What You'll Need Essential: Bachelor's degree in Computer Science, Machine Learning, Statistics or a related field, or equivalent practical experience 3 or more years of experience in ML engineering, NLP or applied data science with hands-on exposure to LLM or agent-based systems Deep practical experience building, chaining and evaluating autonomous agent workflows and frontier LLMs Practical experience building or operating evaluation frameworks, automated scoring or benchmark systems for ML and LLM outputs Strong Python skills and comfort building data pipelines for evaluation datasets Solid understanding of NLP and modern LLM capabilities including prompting techniques, agentic workflows and retrieval Working experience with Google Cloud Platform including Vertex AI and BigQuery, or equivalent AWS or Azure experience Experience with inference optimisation and bridging frontier models into lower-cost internal systems Desirable: Experience with LLM-as-judge techniques, rubric design or human-in-the-loop evaluation programmes Familiarity with agent architectures and the specific failure modes of multi-step and agentic systems Experience operating evaluation systems at scale in a production environment Why Apply? Competitive salary plus equity, with package tailored to UK, NI or European candidates Work on genuinely novel evaluation and agent infrastructure at the competitive core of a scaling AI product Hands-on, intellectually driven team that debates ideas openly and challenges assumptions constructively Async-friendly culture with approximately three syncs per week, designed to protect deep focus and work-life balance Fast-moving startup environment with full operational autonomy and no corporate inertia Belfast based with a globally distributed, elite AI engineering team behind you Interested? For a confidential conversation about this opportunity, connect with Justin Donaldson on LinkedIn or submit your CV via the link below. Skills: ML NLP Python Cloud ML Ops Benefits: Equity TPBN1_NI
Machine Learning Engineer (LLM) employer: Ocho
Join a dynamic and innovative digital consultancy that champions a remote-first work culture, offering exceptional benefits such as a competitive salary, generous annual leave, and a commitment to professional development. With a focus on quality and collaboration, you'll thrive in an environment that values your expertise and provides the autonomy to shape QA processes while working on complex, multi-component systems alongside a talented team in Northern Ireland.