Salary: Β£65,000 - 105,000 per year
Requirements
- We require a Masters or PhD in Computer Science, Data Science, Machine Learning, or a related field, or equivalent practical experience.
- We require experience in data science, machine learning, or applied NLP.
- We require strong hands-on experience with search and retrieval systems, including lexical, vector, and hybrid approaches.
- We require strong hands-on experience with RAG pipelines and LLM-based systems.
- We require strong hands-on experience with evaluation methodologies for ML, IR, and GenAI.
- We require advanced programming skills in Python.
- We require experience with modern ML/NLP frameworks such as PyTorch, Hugging Face, LangChain, LangGraph, or Haystack.
- We require experience working with Databricks or similar distributed data and ML platforms.
- We require a strong understanding of experimentation design and statistical analysis.
- We prefer a PhD in Computer Science, Data Science, Machine Learning, or a related field.
- We prefer experience working with large-scale datasets, including scientific, biomedical, or enterprise data.
- We prefer familiarity with scientific ontologies and metadata standards such as MeSH, UMLS, ORCID, and CrossRef.
- We prefer exposure to production ML systems and MLOps practices.
- We prefer familiarity with data visualization and analytical tools such as Tableau, Power BI, matplotlib, seaborn, or similar.
- We prefer experience with human-in-the-loop evaluation or annotation workflows.
- We prefer publications or demonstrated applied research in IR, NLP, or generative AI.
Responsibilities
- We lead the development and optimization of lexical, vector, and hybrid retrieval systems at scale.
- We help architect and improve RAG pipelines, including retrieval strategies, prompt design, and system orchestration.
- We drive experimentation with embeddings, re-ranking models, and retrieval architectures to improve relevance and user outcomes.
- We partner with engineering to ensure robust, scalable, and production-ready implementations.
- We define and evolve evaluation strategies for search and generative AI systems across our products.
- We design robust frameworks for IR evaluation, including NDCG, recall, and ranking quality.
- We design robust frameworks for GenAI evaluation, including grounding, faithfulness, and hallucination detection.
- We contribute to the development of evaluation datasets, gold standards, and annotation strategies.
- We guide and review experimental design, including offline evaluation and A/B testing, to ensure statistical rigor and validity.
- We contribute to responsible AI practices, including bias, fairness, and risk evaluation.
- We apply and adapt state-of-the-art techniques in NLP, embeddings, and generative AI to production use cases.
- We evaluate and integrate emerging technologies into our roadmap.
- We contribute to knowledge graph and semantic enrichment efforts that support retrieval systems.
- We collaborate with domain experts, ontology engineers, and biomedical informaticians to integrate scientific taxonomies, citation networks, and clinical ontologies into retrieval systems.
- We incorporate structured data, including datasets, chemical entities, genes, drugs, clinical trials, and patient outcomes, into AI-powered discovery pipelines.
- We advance our knowledge graph and metadata integration strategy to enable more context-aware retrieval.
- We apply cutting-edge research in information retrieval, NLP, embeddings, and generative AI to evolve our discovery and evaluation stack.
- We work closely with product, engineering, and domain experts to define and deliver impactful solutions.
- We communicate findings and recommendations clearly to both technical and non-technical stakeholders.
- We take ownership of projects from problem definition through experimentation and deployment.
Technologies
- AI
- Architect
- Databricks
- Support
- LLM
- Machine Learning
- MLOps
- Power BI
- PyTorch
- Python
- RAG
- Tableau
More
We are Elsevier, a global leader in information and analytics, helping researchers, clinicians, and life sciences professionals advance discovery and improve health outcomes through trusted content, data, and analytics. Our Search & AI Evaluation team sits within the Platform Data Science organization and advances enterprise-scale search, retrieval, and evaluation capabilities across our global products, including Scopus AI, LeapSpace, ClinicalKey AI, PharmaPendium, and next-generation life sciences platforms. We offer a culture of innovation, collaboration, and excellence, with a focus on healthy work/life balance, flexible working hours, wellbeing initiatives, shared parental leave, study assistance, sabbaticals, and a wide range of country-specific benefits. The role is full time and is based in London Wall or Amsterdam; if performed in Amsterdam, the base pay range is 53,800 - 89,900.
last updated 36 week of 2026
#J-18808-Ljbffr
Senior Data Scientist - London employer: Elsevier
At Elsevier, we pride ourselves on being an excellent employer, particularly for our Java Software Engineer role in Oxford. Our vibrant work culture fosters collaboration and innovation, while our commitment to employee growth is evident through continuous learning opportunities and flexible working arrangements that promote a healthy work-life balance. Join us to be part of a team that values your contributions and supports your professional journey in the exciting field of scientific knowledge sharing.