Data Engineer - Tech Lead (Databricks, Pyspark)

Data Engineer - Tech Lead (Databricks, Pyspark)

Full-Time 63000 - 77000 £ / year (est.) Home office (partial)
EPAM Systems

At a Glance

  • Tasks: Lead the design and optimisation of scalable data architectures using Azure Databricks and PySpark.
  • Company: Join a forward-thinking tech company in London with a hybrid work culture.
  • Benefits: Competitive salary, flexible working, and opportunities for professional growth.
  • Other info: Collaborative environment with a focus on modern engineering practices and career advancement.
  • Why this job: Shape large-scale data ecosystems and drive innovation in AI-driven solutions.
  • Qualifications: Experience in Azure Databricks, PySpark, and strong leadership skills required.

The predicted salary is between 63000 - 77000 £ per year.

Hybrid in The United Kingdom: London. We're looking for a Senior Data Engineer – Tech Lead (Databricks, PySpark) to join our team in London, UK, in a hybrid working mode.

In this role, you will lead the design, development and optimization of scalable cloud-native data architectures, focusing on Azure Databricks, PySpark and Lakehouse principles. You will work hands-on to deliver performant data solutions for high-volume workloads, ensuring governance, reliability and best practices for enterprise-grade platforms.

As a technical leader, you will define data strategies, drive modernization initiatives and mentor engineers, fostering excellence and innovation throughout the team. This position offers the opportunity to shape large-scale data ecosystems, implement modern engineering practices and enable next-generation analytics and AI-driven solutions.

Responsibilities
  • Lead the architecture, design and build of large-scale data platforms using Azure Databricks and modern cloud technologies.
  • Implement and optimize ETL workflows and streaming pipelines with PySpark and Delta Live Tables following Lakehouse principles.
  • Enhance performance, manage cloud costs and ensure platform reliability for structured streaming workloads.
  • Define data governance, security and quality standards to maintain consistency across the platform.
  • Collaborate with stakeholders to translate complex business requirements into actionable technical solutions.
  • Develop integration approaches using Azure-native services such as Data Factory, Synapse and Blob Storage.
  • Mentor data engineers, promote modern engineering practices and perform technical reviews.
  • Drive adoption of CI/CD, Infrastructure as Code and automated testing in data engineering environments.
  • Implement observability and monitoring using tools like Databricks Workflows and related frameworks.
  • Contribute to AI-driven initiatives by leveraging Databricks ML/MosaicML to integrate Generative AI and LLM-based solutions.
Requirements
  • Bachelor’s or Master’s degree in Computer Science, Software Engineering or related field.
  • Extensive experience designing and implementing production-grade platforms using Azure Databricks.
  • Expertise in PySpark, including advanced optimization, data skew mitigation and query tuning.
  • Strong programming skills in Python with knowledge of modern software design principles.
  • Practical experience with structured streaming, Delta Lake and Delta Live Tables.
  • Proven experience in Lakehouse migration and modernization using open table formats such as Delta Lake or Apache Iceberg.
  • Proficiency with cloud-native services on Azure and knowledge of multi-cloud environments (AWS or GCP).
  • Hands-on experience with CI/CD and Infrastructure as Code tools (Terraform, GitHub Actions, Jenkins).
  • Strong leadership ability to guide teams, define epics/user stories and ensure delivery in agile environments.
  • Excellent communication and stakeholder management skills for both technical and non-technical audiences.
Nice to have
  • Experience operationalizing LLM or Generative AI workflows in Databricks pipelines.
  • Familiarity with frameworks like LangChain, LlamaIndex or Databricks ML/MosaicML.
  • Knowledge of AI governance, security practices and enterprise integration controls.
  • Background in financial trading data or related domains.
  • Official Databricks certifications such as Certified Data Engineer Professional or Apache Spark Developer.

Data Engineer - Tech Lead (Databricks, Pyspark) employer: EPAM Systems

EPAM Systems is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the rapidly evolving field of AI. With a strong focus on employee growth, you will have access to extensive learning opportunities and benefits such as stock purchase plans, all while working remotely or in a hybrid model across Europe. Join us to make a meaningful impact in the energy and utilities sector, driving transformation strategies alongside industry leaders.

EPAM Systems

Contact Details:

EPAM Systems Recruitment Team

We think you need these skills to ace Data Engineer - Tech Lead (Databricks, Pyspark)

SQL
Problem-Solving Skills
Python
Communication Skills
Data Engineering
Data Pipeline Development
API Integration