At a Glance
- Tasks: Lead the design and execution of next-gen data architecture for clinical trials.
- Company: Join Medidata, a leader in digital solutions for smarter treatments and healthier people.
- Benefits: Enjoy competitive pay, generous holidays, and comprehensive health benefits.
- Other info: Hybrid work model with opportunities for professional growth and development.
- Why this job: Make a real impact on life-saving therapies through innovative data engineering.
- Qualifications: Expertise in data engineering, cloud platforms, and AI integration required.
The predicted salary is between 72000 - 88000 £ per year.
This is a hybrid remote/in-office role.
About our Company: Medidata is powering smarter treatments and healthier people through digital solutions to support clinical trials. Celebrating over 25 years of ground-breaking technological innovation across more than 38,000 trials and 12 million patients, Medidata offers industry-leading expertise, analytics-powered insights, and one of the largest clinical trial data sets in the industry. More than 1 million registered users across approximately 2,300 customers trust Medidata's seamless, end-to-end platform to improve patient experiences, accelerate clinical breakthroughs, and bring therapies to market faster. A Dassault Systèmes brand, Medidata is headquartered in New York City and has been recognised as a Leader by Everest Group and IDC.
Our team: At the heart of Medidata's ecosystem, the Data Platform Team powers and connects every application, serving as the engine for enterprise data convergence. Every interaction across our global platform generates critical data—and our mission is to transform that raw information into high-impact insights. Operating on a Data-as-a-Product philosophy, our team builds the foundational infrastructure. This infrastructure includes high-throughput streaming and cloud warehousing to automated data governance. It fuels clinical analytics, AI/ML innovations, and global data sharing.
Why Join Us?
- Direct Impact at Scale: Promote the central data engine behind every Medidata application, directly accelerating clinical trials and life-saving operational outcomes globally.
- Modern Distributed Stack: Build at the intersection of real-time event streaming pipelines, scalable cloud data warehouses, and enterprise-grade automated data security.
- Product-Minded Engineering: Treat data as a first-class product, transforming static databases into high-value, reusable assets for internal AI/ML teams and external partners.
- Fuel Advanced AI/ML: Promote next-generation predictive modelling and clinical analytics by standardising and safeguarding complex healthcare datasets.
What will you do: Reporting to a Director of Engineering, as a Principal Data Engineer / Architect, you will lead the strategic vision and hands-on execution of our next-generation Object-Centric Data Fabric. You will transition traditional application-centric architectures into a centralised semantic layer that seamlessly unifies multi-stream operational data—including Electronic Data Capture (EDC), patient telemetry, and real-world health datasets. In this role, AI augmentation is natively woven into your workflow. It acts as a force multiplier to automate routine mapping, query optimization, and regulatory documentation. This allows you to focus on driving high-impact platform architecture.
- Data Fabric & Lake Architecture: Architect and evolve the enterprise semantic data fabric, converting multi-stream clinical execution datasets into an object-centric model. Design and execute a modern Data Lake strategy centred on Apache Iceberg as the core storage format, ensuring high-performance querying and seamless interoperability with Snowflake and heterogeneous compute engines.
- AI-Accelerated Schema & Pipeline Engineering: Develop and maintain end-to-end multi-stream ingestion pipelines for complex clinical trial schemas. Use AI-driven schema inference and ontology alignment tools to auto-draft mapping artifacts, dramatically reducing integration timelines across different life science datasets.
- High-Throughput Streaming & Backend Services: Build scale, fault-tolerant real-time ingestion pipelines using Kafka, AWS, and Snowflake. Write robust enterprise services in Java or Scala, leveraging AI coding assistants for rapid code generation, refactoring, and performance tuning.
- Technical Strategy & Database Optimization: Promote technical direction and engineering best practices across teams for Change Data Capture (CDC), clustering, data migration, and aggregation. Use AI query-optimization tools to analyse execution plans, auto-tune complex Snowflake/Iceberg SQL workloads, and eliminate performance bottlenecks.
- Intelligent Telemetry & Closed-Loop Reasoning: Integrate automated AI inferencing and reasoning layers directly into data pipelines to detect telemetry anomalies in real-time and automatically map safety signals back to operational trial datasets.
- Quality & AI-Driven Compliance: Lead Test-Driven Development (TDD) and Behavior-Driven Development (BDD) initiatives. Use AI test generators to produce HIPAA/GxP-compliant synthetic clinical trial datasets for automated validation. Leverage AI tools to auto-draft validation artifacts, data lineage manifests, and audit documentation required under GxP, HIPAA, and GDPR standards.
- System Resilience: Troubleshoot complex production issues across distributed data environments and implement resilient, self-healing pipeline architectures.
Key Business Value & Strategic Impact: Your leadership will directly advance life sciences technology. By pairing modern lakehouse architecture (Snowflake, Apache Iceberg, Kafka) with AI-augmented workflows, you will empower our platform to process critical safety signals in hours rather than weeks. This enables adaptive clinical trial execution, guarantees zero-loss data integrity, and significantly accelerates regulatory submission timelines for life-saving therapies.
Qualifications
- Education & Experience: Bachelor's or Master's degree in Computer Science, Data Science, Software Engineering, or equivalent practical experience, alongside proven years of dedicated professional experience in enterprise data engineering and architecture.
- Data Warehousing & Data Lake Mastery: Expert-level mastery of SQL (OLAP/OLTP) and enterprise cloud data platforms (specifically Snowflake), combined with practical experience architecting open data lake structures using Apache Iceberg.
- Software Engineering & Streaming: Demonstrated expertise building production backend services in Java or Scala on AWS, alongside deep familiarity with real-time streaming architectures (e.g., Apache Kafka) and modern data engineering design patterns.
- AI-Augmented Engineering Proficiency: Active, practical experience integrating AI developer tools (e.g., GitHub Copilot, Cursor) into daily workflows to accelerate SQL query generation, code refactoring, automated testing, and technical documentation drafting.
- Engineering Rigor & Methodologies: Proven command of Git revision control, CI/CD pipeline automation, and TDD/BDD practices, augmented by AI-driven test case generation and quality checks.
- Domain & Regulatory Awareness: Strong foundational understanding of clinical trial data workflows (e.g., EDC architectures), healthcare data models, and life science compliance standards (GxP, HIPAA, GDPR), with the ability to apply AI/ML tools for schema mapping and zero-loss data integrity verification.
Base pay is one part of the Total Rewards that Medidata provides to compensate and recognise employees for their work. Most sales positions are eligible for a commission on the terms of applicable plan documents, and many of Medidata's non-sales positions are eligible for annual bonuses. Medidata believes that benefits should connect you to the support you need when it matters most and provides benefits, including medical, dental, life and disability insurance; a generous pension; and 25+ paid holidays per year.
We will accept applications on an ongoing basis until we fill the position.
Principal Data Engineer in London employer: MEDIDATA
Medidata is an exceptional employer, offering a dynamic hybrid work environment that fosters innovation and collaboration in the healthcare technology sector. With a commitment to employee growth, Medidata provides extensive benefits including comprehensive health insurance, a generous pension plan, and over 25 paid holidays annually, ensuring a supportive work culture that values work-life balance. Join us to make a direct impact on clinical trials and patient outcomes while working with cutting-edge technologies in a company recognised for its leadership in the industry.
StudySmarter Expert Advice🤫
We think this is how you could land Principal Data Engineer in London
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like MEDIDATA!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Principal Data Engineer at MEDIDATA.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like MEDIDATA.
✨Apply Directly through Our Website
When you find a suitable opening like Principal Data Engineer at MEDIDATA, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Principal Data Engineer in London
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at MEDIDATA, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at MEDIDATA. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at MEDIDATA
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at MEDIDATA!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.