Data Ingestion Engineer in London

Data Ingestion Engineer in London

London Full-Time 80000 - 98000 £ / year (est.) No working from home possible
Cognitive Group | Part of the Focus Cloud Group

At a Glance

  • Tasks: Solve practical, high-impact data problems and optimise ingestion pipelines.
  • Company: Dynamic tech company focused on innovative data solutions.
  • Benefits: Competitive pay, hybrid work model, and opportunities for professional growth.
  • Other info: Collaborative environment with a focus on innovation and career development.
  • Why this job: Make a real impact on large-scale data systems and enhance your engineering skills.
  • Qualifications: Experience with Apache Spark and large-scale data processing.

The predicted salary is between 80000 - 98000 £ per year.

This is a hands‑on contract role for engineers who enjoy solving practical, high-impact problems at scale.

You will help keep our ingestion pipelines running smoothly, unblock critical data flows, and improve the reliability of systems that support annotation, data science, model training and evaluation.

6-month Contract role - Inside IR35

Responsibilities

  • Debug and fix failing or blocked ingestion pipelines.
  • Investigate issues caused by corrupt, malformed or unexpected data.
  • Help make our pipelines more resilient so individual bad data segments do not block wider workflows.
  • Improve how we handle varied data formats from partners, suppliers and third‑party sources.
  • Support orchestration across multi‑step ingestion workflows, including dependencies, retries and queue management.
  • Optimise Spark jobs and data processing pipelines for throughput and compute efficiency.
  • Reduce operational toil around failed jobs, stalled pipelines and manual interventions.
  • Work on high‑volume batch processing systems where throughput, reliability and cost all matter.
  • Help prioritise and unblock important datasets for downstream annotation, data science and model training teams.
  • Partner with permanent engineers to keep critical ingestion work moving while longer‑term platform improvements are developed.
  • Contribute pragmatic improvements that make the system easier to operate, scale and trust.

Qualifications

You are an experienced data engineer, platform engineer or distributed systems engineer who enjoys working on large‑scale production data systems.

  • Required Skills
  • Strong production experience with Apache Spark.
  • Experience building, debugging or operating large‑scale data ingestion, ETL or data processing pipelines.
  • Experience with distributed data processing systems.
  • Ability to optimise jobs for throughput, compute efficiency and reliability.
  • Comfort working with messy, corrupt, incomplete or inconsistent data.
  • Understanding of orchestration across multi‑step pipelines and downstream dependencies.
  • Ability to work independently in a fast‑moving, highly technical environment.
  • A practical, delivery‑focused mindset and the ability to ramp quickly.
  • Experience working at significant data scale, ideally PB‑scale or similarly high‑throughput environments.
  • Preferred Skills
  • Experience in one or more of the following areas would be a strong advantage:
  • Airflow, Flyte, Databricks workflows or similar orchestration tooling.
  • Scala or Java, especially in Spark‑based environments.
  • Queue‑based processing, retry handling and priority data workflows.
  • High‑throughput batch data processing systems.
  • Production systems with many data producers, consumers or external data sources.
  • Handling third‑party, partner or supplier data with inconsistent formats and quality issues.
  • Automotive, robotics, autonomy, mapping, ML data platforms or embodied AI environments.
  • Cost optimisation for compute‑ and storage‑heavy data platforms.
  • High‑performance engineering experience from domains such as trading, where it includes relevant distributed systems or throughput‑focused work.
  • Pay range and compensation package

This is a full‑time contract role based in our office in London.

We operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home.

#J-18808-Ljbffr

Data Ingestion Engineer in London employer: Cognitive Group | Part of the Focus Cloud Group

As a leading Microsoft consultancy, we pride ourselves on fostering a dynamic work culture that champions innovation and continuous learning. Our Technical Copilot Trainer role offers not only the opportunity to engage with cutting-edge AI technologies but also provides a clear pathway for career advancement into AI Consulting or Solution Architecture. With a hybrid working model based in London, employees benefit from a collaborative environment while enjoying the flexibility to balance their professional and personal lives.

Cognitive Group | Part of the Focus Cloud Group

Contact Details:

Cognitive Group | Part of the Focus Cloud Group Recruitment Team

We think you need these skills to ace Data Ingestion Engineer in London

Apache Spark
Data Ingestion
ETL
Data Processing Pipelines
Distributed Data Processing Systems
Job Optimisation
Data Quality Management