About the Role
We are seeking a driven, analytical, and detail-oriented Data Engineer with 1 to 3 years of hands-on professional experience to join our growing Data Platform team. In this role, you will bridge the gap between raw data generation and actionable business insights. You will play a pivotal part in architecting, developing, testing, and maintaining scalable data architectures, robust pipelines, and distributed compute platforms that fuel our enterprise intelligence.
This is an exceptional opportunity for an early-to-mid career engineer looking to deepen their technical skills. You will work closely with seasoned Staff Engineers, Data Scientists, Product Managers, and Business Analysts to transform complex unstructured and structured datasets into highly reliable, analysis-ready data pipelines.
Core Technical Requirement: The 13 Key Proficiencies
To succeed in this position, candidates are expected to demonstrate working knowledge and practical experience across our Core 13 Competencies:
- SQL: Advanced querying, data modeling, window functions, and query performance tuning.
- Python: Object-oriented design, scripting, automation, packaging, and data manipulation.
- Apache Spark: Distributed computing, memory optimization, and PySpark framework usage.
- Data Modeling: Star schema, Snowflake schema, Data Vault, and dimensional modeling principles.
- Cloud Platforms: Working knowledge of modern cloud ecosystems (AWS, GCP, or Microsoft Azure).
- Data Warehousing: Direct experience managing and scaling warehouses (Snowflake, BigQuery, or Redshift).
- Workflow Orchestration: DAG creation and dependency tracking using Apache Airflow, Prefect, or Dagster.
- ETL/ELT Architecture: Designing idempotent data ingestion and transformation lifecycles.
- Version Control (Git): Branching strategies, code reviews, pull requests, and collaborative workflows.
- Containerization: Dockerizing applications, services, and local data development environments.
- CI/CD Frameworks: Automated pipeline deployment, continuous integration, and test automation.
- Data Quality & Governance: Implementing testing frameworks (e.g., Great Expectations, dbt tests) and schema validations.
- Data Streaming Basics: Familiarity with message brokers and event streaming concepts (Apache Kafka or AWS Kinesis).
Detailed Responsibilities
As a Data Engineer on our team, your daily mission involves engineering reliability, precision, and efficiency into our data operations. Your core responsibilities include:
1. Pipeline Design, Development, and Maintenance
- Design, implement, and maintain highly scalable, fault-tolerant batch and near-real-time ELT/ETL pipelines that ingest billions of events weekly from transactional databases, third-party APIs, and event logs.
- Write clean, modular, maintainable, and well-documented Python and SQL code following company-wide standards and best practices.
- Schedule, monitor, and troubleshoot distributed jobs via modern workflow orchestrators (Apache Airflow), mitigating DAG failures and ensuring strict delivery against Service Level Agreements (SLAs).
2. Data Modeling and Storage Architecture
- Collaborate with analytics engineers to model complex transactional domains into high-performance analytical datasets leveraging dimensional modeling concepts (fact and dimension tables).
- Manage partitioned, clustered, and optimized data lakehouse formats (Delta Lake, Apache Iceberg, or Parquet) to lower cloud storage costs and accelerate computational runtimes.
- Continuously audit warehouse performance, resolving bottleneck queries, implementing index strategies, and managing cluster auto-scaling rules.
3. Data Quality, Reliability, and Observability
- Embed automatic validation frameworks to evaluate data freshness, completeness, schema drift, and uniqueness at every stage of transformation.
- Build proactive alerting mechanisms to notify cross-functional stakeholders of anomalies, data drops, or upstream API breakages before they impact reporting layers.
- Conduct root cause analysis (RCA) on data-related system outages and implement long-term structural remediations.
4. Cross-Functional Collaboration & Support
- Interface directly with Data Scientists and Machine Learning Engineers to curate engineered features, feature stores, and clean datasets for model training and deployment.
- Empower business stakeholders by maintaining accurate data dictionary definitions, metadata registries, and architecture diagrams.
- Participate in agile sprints, daily standups, backlog grooming, and code reviews to refine engineering velocity and raise software quality.
Qualifications and Candidate Profile
We are looking for candidates who demonstrate a balance between analytical problem-solving and software development fundamentals.
Required Qualifications:
- Professional Experience: 1 to 3 years of dedicated, full-time work experience in a Data Engineering, Software Engineering, or Database Administration capacity.
- Educational Background: Bachelor’s or Master's degree in Computer Science, Software Engineering, Information Systems, Applied Mathematics, or equivalent demonstrable experience.
- Demonstrated Proficiency: Verifiable hands-on application of the 13 core competencies outlined above within enterprise or high-growth production environments.
- Software Engineering Fundamentals: A solid foundation in algorithms, data structures, complexity analysis (Big-O notation), and test-driven development (TDD).
- Problem Solving: Strong analytical and troubleshooting mindset with the ability to isolate pipeline faults across network, compute, and application layers.
Preferred Qualifications:
- Experience utilizing dbt (data build tool) for modern warehouse-centric transformations.
- Exposure to Infrastructure as Code (IaC) tooling, such as Terraform or AWS CloudFormation.
- Prior contributions to enterprise lakehouse migrations or distributed compute optimization initiatives.
- Relevant certifications: AWS Certified Data Engineer, GCP Professional Data Engineer, or Databricks Certified Associate Developer.
What We Offer
We believe in investing in our team members, fostering continuous learning, and offering a collaborative environment where data drives decision-making at every level.
- Competitive Compensation: Market-leading base salary with annual performance bonuses and equity incentives.
- Health & Wellness: Comprehensive health, dental, and vision insurance with flexible spending accounts (FSA/HSA).
- Work Flexibility: Hybrid working options with home-office setup stipends and flexible hours.
- Continuous Growth: Dedicated annual training and development budget for technical conferences, cloud certifications, and coursework.
- Time Off: Generous paid time off (PTO), company holidays, paid parental leave, and regular company recharge days.
- Retirement Savings: 401(k) retirement plan with immediate employer matching.
How to Apply
If you meet the requirements, possess working mastery across the 13 proficiencies, have 1 to 3 years of production-level experience, and are eager to tackle complex data challenges at scale, we want to hear from you. Please submit your application below with your resume and a link to your GitHub profile or relevant portfolio projects.
Data Engineer in London employer: Recruiterflow Staging
As a Senior Director Manager at our company, you will thrive in a dynamic and supportive work environment that prioritises employee growth and development. We offer competitive benefits, a collaborative culture, and opportunities for professional advancement, all set in a vibrant location that fosters innovation and creativity. Join us to lead a dedicated team and make a meaningful impact while enjoying a fulfilling career.