London based, 1-2 office days per week, with monthly travel to the UAE. Willingness to relocate to the UAE is optional.
Company & role
This role sits with a global IT solutions provider building a new AI delivery capability for banking, insurance and fintech clients. AI is already central to what they do. This team build is about scaling that into production grade solutions their financial services customers can rely on.
You will own the quality gate across every prototype and production release, making sure nothing ships without measured, documented evidence that it is defensible to client compliance and model risk teams. In short, you turn "it appears to work" into auditable quality. That means both solid full stack testing and, crucially, the harder problem of evaluating non deterministic AI output for accuracy, safety and reliability.
This is a genuinely important hire in a regulated setting, where the evidence you produce is what lets the business stand behind what it has built.
Why This Role Stands Out
- Quality here is not an afterthought bolted on at the end. It is the thing that lets the business ship AI to banks with confidence, and you own it.
- You get to work at the frontier of AI quality, building evaluation frameworks, golden datasets and red team tests for LLM and agentic systems, which is a genuinely scarce and growing skill set. You will shape how quality and release gates work across a new capability rather than inheriting someone else's process, sit close to real model risk and compliance problems, and work with a modern AI testing stack. Add regular time in the UAE and this is a role that stretches well beyond conventional QA.
Key Responsibilities
- Own the quality gate across prototypes and production releases, ensuring nothing ships without documented, defensible evidence
- Design test strategy and plans covering happy paths, edge cases, integration, performance and security
- Build and maintain automated test suites across unit, integration and end to end, running in CI and CD to enable fast, confident releases
- Validate AI products before client release using golden datasets, real world scenarios, regression suites and agreed thresholds for accuracy, relevance, safety, latency and cost
- Evaluate non deterministic LLM and agent output using semantic evaluation, LLM as judge scoring, human review for high risk cases and repeat testing to detect instability and drift
- Test RAG systems for retrieval quality, answer faithfulness, citation accuracy and fallback behaviour when knowledge is missing or ambiguous
- Run adversarial and red team testing for prompt injection, jailbreaks, data leakage, insecure tool use and biased output before deployment
- Own client release gates, blocking release where critical risks remain open, evaluation scores fall below threshold, or monitoring and rollback plans are incomplete
- Manage defects end to end, from clear reproduction through triage, root cause analysis and driving improvements back into test plans and CI gates
- Ensure production observability is in place for prompts, responses, traces, evaluation scores, latency, cost and drift
Ideal Experience
- 4 or more years in QA, testing or quality engineering
- 2 or more years of test automation using Playwright, Cypress, Selenium or similar, writing maintainable test code
- Full stack testing experience across user interface, API, database and integration scenarios, with a real understanding of how components interact
- Agile testing experience within short sprints and continuous integration
- A strong analytical mindset, comfortable with root cause analysis, metrics and designing effective test plans
- Financial services, banking, insurance or fintech domain experience (mandatory)
- AI and machine learning testing, including evaluating agent and LLM behaviour and defining quality metrics for non deterministic systems
- Experience with a modern AI testing stack such as Promptfoo, DeepEval, RAGAS, LangSmith, Braintrust or Langfuse
- Performance and load testing with k6 or JMeter
- Security testing knowledge including OWASP and penetration testing basics
- BDD frameworks such as Cucumber and Gherkin
- Understanding of UAE compliance and regulated environment testing
#J-18808-Ljbffr