Director AI Systems Reliability, Testing & Performance - Novartis

Director AI Systems Reliability, Testing & Performance - Novartis

Full-Time No working from home possible
O

Summary

Director: AI Systems Reliability, Testing & Performance #LI-Hybrid Location: London Novartis is unable to offer relocation support for this role: please only apply if this location is accessible for you. The Director, AI Systems Reliability, Testing & Performance is a senior technical AI leadership role within Data Science & AI, responsible for defining how AI systems across Novartis Development are evaluated, tested, validated, benchmarked, monitored, and continuously improved throughout their lifecycle. This role owns the technical evidence required to determine whether AI systems are reliable, robust, secure, and fit-for-use across Novartis Development. The role defines common approaches for AI evaluation, benchmarking, testing, validation, monitoring, and production-readiness across agentic AI systems, predictive models, retrieval systems, digital twins, and other AI capabilities. The Director provides technical leadership in AI evaluation science, reliability engineering, validation, adversarial testing, and performance assessment. Key areas of focus include model and agent evaluation, benchmarking, failure-mode analysis, drift detection, digital twin validation, AI red teaming, observability, traceability, and technical evidence generation supporting regulated and business-critical AI systems. This is not a Governance, Product Management, PMO, or infrastructure operations role. Governance owns policies, risk frameworks, and approval processes. Product teams own roadmaps, adoption, and value realization. DDIT and engineering teams own platforms, infrastructure, and operational services. Success means Development AI systems are supported by objective evidence demonstrating how they perform, where they fail, and whether they remain fit-for-use over time.

About the Role

Major Accountabilities

AI Evaluation & Benchmarking

  • Define evaluation methodologies for AI systems, models, agents, digital twins, and simulation environments across Development.
  • Establish benchmark suites, evaluation datasets, and testing harnesses used across AI initiatives.
  • Define objective measures for quality, reliability, robustness, grounding, agent effectiveness, and task success.
  • Ensure evaluation approaches remain scientifically rigorous, reproducible, and comparable across AI systems.
  • Build a common evidence framework for assessing AI capabilities, limitations, and fitness-for-use.

AI Testing & Validation

  • Establish approaches for hallucination testing, failure-mode analysis, robustness testing, and behavioral validation.
  • Define validation methodologies for agentic systems, digital twins, simulation environments, and human-in-the-loop workflows.
  • Develop production readiness criteria for AI systems operating in Development environments.

#J-18808-Ljbffr

Director AI Systems Reliability, Testing & Performance - Novartis employer: OpenTalent

At TwinStream, we pride ourselves on being an exceptional employer, offering a dynamic work culture that fosters collaboration and innovation. Our hybrid working model in Cheltenham allows for flexibility while providing ample opportunities for professional growth and development within the tech sector. With a focus on meaningful projects that support government organisations, employees can find purpose in their work while enjoying competitive benefits and a supportive team environment.

O

Contact Details:

OpenTalent Recruitment Team