Are you interested in working with data and analytics to solve problems?Are you interested in bringing your GenAI, ML and NLP expertise to projects?About our TeamData Science Life Sciences is a diverse team focusing on GenAI, ML, NLP. We mainly develop best-in-class enrichment pipelines for Elsevier’s life science .com products such as Reaxys, Embase and Pharmapendium.About the RoleAs a Senior Data Scientist, you will play a pivotal role in the development and deployment of cutting-edge Gen AI models and solutions. You will be responsible for building, testing, and maintaining our Gen AI, RAG and NLP solutionsYou will work throughout the whole life cycle of data science projects: design, implementation, production and beyond. You will deliver efficient and production-ready Python code. You will collaborate closely with developers to deploy and productionize our data science pipelines and with subject matter experts in biology and chemistry domains to validate the output.This role requires a strong foundation in Natural Language Processing (NLP), Machine Learning, Transformer models and Generative AI, as well as proficiency in Python.ResponsibilitiesData collection, data analysis, model development, defining quality metrics, quality assessment of models and regular presentations to stakeholders.Creating production-ready Python packages for each component of data science pipelines (such as pre-processing and model inference) and their deployment together with software engineering teamOptimizing and customizing Retrieval Augmented Generation (RAG) pipelines to meet specific project requirements that involve content ingestion, machine translation, and contextualized information retrievalIngesting, preprocessing, and transforming large-scale multilingual data to ensure high-quality inputs for downstream models.Building AI agentic models integrated with RAG pipelines.Conducting rigorous testing and evaluation of AI models to ensure high performance and reliability.Integrating data science components and performing end-to-end quality assessments.Maintaining robustness of data science pipelines against model drift and ensuring consistent output quality.Establishing reporting processes for pipeline performance and developing automated re-training strategies for existing pipelines.Collaborating with cross-functional teams to integrate AI solutions into existing products and services.Leading and managing projects with a team of data scientists and independently executing the entire small-scale projectsMentoring junior data scientists and fostering a knowledge-sharing culture within the team.Staying up-to-date with the latest advancements in AI, machine learning, and NLP technologies.RequirementsMaster’s or Ph.D. in Computer Science, Data Science, Artificial Intelligence, or a related field.5+ years of relevant applied experience in data science, with a focus on Generative AI, NLP, and machine learning.Proficiency in Python for data analysis, model development, and deployment.Strong experience with transformer modelsProficiency in Generative AI technologies, including utilizing LLMs via API access, LLM evaluation tools, and prompt engineering.Knowledge of various RAG pipelines and their practical implementation.Experience building Agentic RAG systems is strong requirement.Experience with AI agent management frameworks such as LangChain, or similar tools.Experience with advanced algorithms in deep learning, neural networks, reinforcement learning, and
Senior Data Scientist I in Kennington employer: Elsevier
At Elsevier, we pride ourselves on being an excellent employer, particularly for our Java Software Engineer role in Oxford. Our vibrant work culture fosters collaboration and innovation, while our commitment to employee growth is evident through continuous learning opportunities and flexible working arrangements that promote a healthy work-life balance. Join us to be part of a team that values your contributions and supports your professional journey in the exciting field of scientific knowledge sharing.