At a Glance
- Tasks: Lead data engineering architecture and standards for groundbreaking AI projects in healthcare.
- Company: Join Boehringer Ingelheim, a top employer dedicated to innovative therapies and exceptional workplace culture.
- Benefits: Enjoy a hybrid work model, competitive salary, and opportunities for professional growth.
- Other info: Collaborate with a dynamic team and shape the future of healthcare technology.
- Why this job: Make a real impact on disease understanding and therapeutic development through cutting-edge data engineering.
- Qualifications: PhD or MSc in STEM with extensive data engineering experience and strong biomedical knowledge.
The predicted salary is between 72000 - 88000 Β£ per year.
Most diseases are still poorly understood at a biological level. Despite decades of research, the causal mechanisms driving many conditions remain unclear, limiting our ability to identify the right targets, design the right interventions and bring the right medicines to patients. The AI Accelerator exists to change that. Based in London and sitting within Computational Innovation, a global organisation spanning computational biology, human genetics, data excellence and AI, the Accelerator's mission is to build production-quality AI capabilities that deepen our understanding of disease biology and increase probability of success.
We do this by applying neural-based methods across the biomedical data landscape to integrate heterogeneous, multimodal data sources, infer biological relationships and embed causal thinking into what we build. The goal is not just to predict but to explain and understand why disease occurs. It could be electronic health records and medical imaging to support patient segmentation. It could be 'omics data to identify novel therapeutic targets. It could be predicting transcriptional change for a given disease-causing variant. It could be simulating the effect of modulating a target of interest.
None of this is possible without data. The quality, accessibility and engineering of biomedical data ultimately determine the quality of the models built upon them. The Data Engineering function provides the foundation on which the AI Accelerator operates by integrating diverse clinical, biological and real-world datasets into trustworthy, AI-ready assets that can support large-scale foundation model development and downstream scientific discovery.
Key Responsibilities
- Set the technical direction, strategy and roadmap for data engineering, aligned to AI Accelerator priorities.
- Own the data engineering architecture for the AI Accelerator, defining and evolving a layered architecture, feature and embedding provisioning patterns, and the harmonised multimodal data foundation that underpins model development.
- Establish data engineering standards and engineering practices, including data quality controls, testing, CI/CD for data, reproducibility standards, data contracts, metadata management and dataset versioning.
- Deliver hands-on leadership on the most difficult and highest-impact data engineering challenges, including integrating novel and complex modalities such as genomics, transcriptomics, imaging, clinical and real-world datasets.
- Partner with the Data Excellence community and IT to align on governance, ontologies, metadata harmonisation, stewardship and shared enterprise data foundations.
- Establish ways of working and coach other team members, onboarding, mentoring and technically leading data engineers as the team scales, while acting as the senior escalation point for complex data engineering challenges.
Requirements
- PhD or MSc and equivalent experience in a STEM subject.
- Extensive experience operating at the senior staff level within a data engineering function, including defining architecture, standards and technical direction, along with mentoring data engineers and establishing strong engineering principles within a growing team.
- Deep expertise in large-scale data engineering, including distributed processing frameworks, pipeline orchestration, cloud data platforms and data integration at scale, with experience integrating internal and third-party datasets and working effectively with external data providers and technology partners.
- Strong understanding of biomedical and healthcare data domains, such as genomics, transcriptomics, multi-omics, imaging, clinical data, electronic health records or real-world data, alongside an understanding of machine learning data requirements, including versioning, reproducibility and tensor-based workflows for large-scale AI systems.
- Experience implementing data governance in practice, including metadata management, lineage, provenance, ontologies, cataloguing and FAIR principles and familiarity with Trusted Research Environments (TREs) and controlled-access research data environments.
- Strong collaboration and influencing skills across technical and non-technical stakeholders, with the ability to communicate complex technical concepts clearly.
This is a hybrid role with approximately 4 days a week in the office.
Why This Is A Great Place To Work
Boehringer Ingelheim has been recognised as a Top Employer in the UK, demonstrating our commitment to building an exceptional workplace through strong people practices and supportive HR policies.
Senior Staff Data Engineer employer: Boehringer Ingelheim GmbH
Boehringer Ingelheim is an exceptional employer, recognised as a Top Employer in the UK, offering a supportive work culture that prioritises employee well-being and professional growth. As a Senior ML Engineer in London, you will be at the forefront of biomedical AI, collaborating with leading scientists and engineers to make impactful contributions to human health while enjoying a hybrid work model that promotes work-life balance.