Staff Data Engineer (Emerald)

Staff Data Engineer (Emerald)

Full-Time 63000 - 77000 £ / year (est.) Working from home possible
H

At a Glance

  • Tasks: Lead a team to design and optimise scalable data systems for healthcare datasets.
  • Company: Join H1, a leader in healthcare data solutions with a focus on innovation.
  • Benefits: Enjoy flexible hours, stock options, unlimited PTO, and health insurance.
  • Other info: Collaborative environment with opportunities for mentorship and career growth.
  • Why this job: Make a real impact in healthcare by improving data accuracy and efficiency.
  • Qualifications: 8+ years in data engineering with expertise in Spark, AWS, and Python.

The predicted salary is between 63000 - 77000 £ per year.

As a Staff Data Engineer on the Emerald team, you will play a critical role in shaping the architecture, scalability, and technical direction of H1’s healthcare entity resolution platform. EMERALD is responsible for linking large-scale external healthcare datasets, including PubMed, clinical trials, conferences, ct.gov, and web-collected data to H1’s canonical physician and organization profiles. This role sits at the intersection of distributed data engineering, entity matching, identity resolution, and large-scale healthcare data processing.

You will lead a small team of engineers while remaining deeply hands-on technically, owning the systems and pipelines powering automatching, grouping logic, identity mapping, deduplication, and enrichment workflows processing tens of millions of records. You will partner closely with Product, AI/ML, Analytics, and Engineering teams to improve platform accuracy, scalability, reliability, and operational efficiency across one of H1’s most critical data platforms.

  • Lead the design, optimization, and scalability of distributed Spark/PySpark pipelines powering entity resolution and large-scale healthcare data processing.
  • Own systems supporting automatching, identity mapping, grouping logic, deduplication, enrichment, and auto-approval workflows across healthcare provider and organization datasets.
  • Build and maintain scalable processing frameworks for PubMed, clinical trial, ct.gov, conference, and other healthcare data sources.
  • Drive infrastructure optimization initiatives focused on improving throughput, runtime, observability, and cloud compute cost efficiency.
  • Partner closely with AI/ML teams to integrate matching and resolution models into EMERALD and improve matching precision and recall.
  • Lead complex technical initiatives from architecture and design through deployment, monitoring, and long-term production support.
  • Serve as a technical leader and mentor across the team through code reviews, technical guidance, and engineering best practices.
  • Collaborate directly with Product and business stakeholders to align technical solutions with operational and customer needs.
  • Support production operations, incident response, troubleshooting, and ongoing platform reliability.

Benefits:

  • Flexible work hours
  • Commuter benefits
  • Stock options
  • Computer setup
  • Work from home opportunities
  • Health & life insurance
  • Retirement options
  • Unlimited PTO
  • Flex Give & Flex Spend
  • Impactful BRGs

Experience:

  • Building scalable ETL/ELT frameworks across both batch and streaming architectures.
  • Extensive experience with Apache Spark and AWS-based big data technologies including EMR, S3, and distributed compute environments.
  • Optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiency.
  • Working with containerization and infrastructure technologies such as Docker, Kubernetes, and Terraform.
  • Working with relational or distributed databases such as PostgreSQL or Redshift.
  • Strong hands-on engineering expertise across distributed computing, large-scale data processing, and infrastructure optimization.
  • Strong grasp of software engineering fundamentals including distributed systems, data structures, concurrency, and system design.
  • Experience working with healthcare, life sciences, Real World Evidence (RWE), or large-scale healthcare datasets is strongly preferred.
  • Strong coding experience in Python (PySpark), Scala, Java, or equivalent languages used for distributed processing systems.
  • Deep expertise with distributed data processing frameworks such as Apache Spark and Hadoop, particularly within AWS environments.
  • Strong communication and collaboration skills across both technical and non-technical stakeholders.
  • Improving performance, scalability, observability, and infrastructure efficiency within distributed systems.
  • Experience with streaming and event-driven architectures using technologies such as Kafka or Spark Streaming.
  • Experience with entity resolution, identity mapping, automatching, deduplication, or large-scale matching systems is strongly preferred.
  • Proven ability to operate effectively within highly scalable, production-grade distributed systems.
  • Familiarity with modern development and infrastructure tooling including Git, CI/CD pipelines, Docker, Kubernetes, Terraform, Argo, Hudi, and JIRA.
  • 8+ years of experience building and maintaining large-scale distributed data systems and pipelines.
  • Experience with streaming technologies such as Kafka, Spark Streaming, or KSQL.
  • Performing root cause analysis across large-scale distributed systems and complex data pipelines.
  • Strong proficiency in Python (PySpark), Scala, Java, or other modern programming languages used for large-scale distributed processing.
  • Demonstrated technical leadership experience mentoring engineers and driving complex technical initiatives.
  • Experience with orchestration and lakehouse technologies such as Argo and Hudi or comparable platforms.

You are an experienced data engineer with deep expertise building and optimizing distributed data systems in cloud-native environments. You thrive solving complex scalability and performance challenges across high-volume data processing systems and enjoy operating in highly technical, fast-paced engineering environments. Ability to write clean, maintainable, modular, and production-grade code. Strong understanding of distributed file formats including Apache Parquet and Apache AVRO.

Staff Data Engineer (Emerald) employer: h1

At H1, we pride ourselves on being an exceptional employer that champions health equity and innovation in the healthcare sector. Our inclusive work culture fosters collaboration and empowers employees to grow through meaningful engagement with leading biotech and life sciences companies. With generous benefits, flexible work arrangements, and a commitment to diversity, H1 offers a rewarding environment for those passionate about making a difference in healthcare.

H

Contact Details:

h1 Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Staff Data Engineer (Emerald)

Get Involved in Data Science Meetups

Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like h1!

Show Off Your Projects

Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Staff Data Engineer (Emerald) at h1.

Leverage Professional Networks

Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like h1.

Apply Directly through Our Website

When you find a suitable opening like Staff Data Engineer (Emerald) at h1, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!

We think you need these skills to ace Staff Data Engineer (Emerald)

Apache Spark
PySpark
AWS (EMR, S3)
ETL/ELT frameworks
Docker
Kubernetes
Terraform

Some tips for your application 🫡

Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!

Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!

Craft a Tailored Cover Letter:For a full-time role at h1, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.

Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at h1. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!

How to prepare for a job interview at h1

Brush Up on Your Statistics

For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!

Showcase Your Projects

Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!

Get Comfortable with Python and R

Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at h1!

Prepare for Case Studies

Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.