Data Engineer - Tech Lead (Databricks, Pyspark)

Data Engineer - Tech Lead (Databricks, Pyspark)

Full-Time 63000 - 77000 £ / year (est.) Home office (partial)
EPAM Systems, Inc.

At a Glance

  • Tasks: Lead the design and optimisation of scalable data architectures using Azure Databricks and PySpark.
  • Company: Join a forward-thinking tech company in London with a hybrid work culture.
  • Benefits: Enjoy competitive salary, private medical insurance, and perks like free lunches and on-site massages.
  • Other info: Great opportunities for learning, development, and career growth.
  • Why this job: Shape large-scale data ecosystems and drive innovation in AI-driven solutions.
  • Qualifications: Experience in Azure Databricks, PySpark, and strong leadership skills required.

The predicted salary is between 63000 - 77000 £ per year.

We're looking for a Senior Data Engineer – Tech Lead (Databricks, PySpark) to join our team in London, UK, in a hybrid working mode. In this role, you will lead the design, development and optimization of scalable cloud-native data architectures, focusing on Azure Databricks, PySpark and Lakehouse principles. You will work hands-on to deliver performant data solutions for high-volume workloads, ensuring governance, reliability and best practices for enterprise-grade platforms.

As a technical leader, you will define data strategies, drive modernization initiatives and mentor engineers, fostering excellence and innovation throughout the team. This position offers the opportunity to shape large-scale data ecosystems, implement modern engineering practices and enable next-generation analytics and AI-driven solutions.

Responsibilities
  • Lead the architecture, design and build of large-scale data platforms using Azure Databricks and modern cloud technologies.
  • Implement and optimize ETL workflows and streaming pipelines with PySpark and Delta Live Tables following Lakehouse principles.
  • Enhance performance, manage cloud costs and ensure platform reliability for structured streaming workloads.
  • Define data governance, security and quality standards to maintain consistency across the platform.
  • Collaborate with stakeholders to translate complex business requirements into actionable technical solutions.
  • Develop integration approaches using Azure-native services such as Data Factory, Synapse and Blob Storage.
  • Mentor data engineers, promote modern engineering practices and perform technical reviews.
  • Drive adoption of CI/CD, Infrastructure as Code and automated testing in data engineering environments.
  • Implement observability and monitoring using tools like Databricks Workflows and related frameworks.
  • Contribute to AI-driven initiatives by leveraging Databricks ML/MosaicML to integrate Generative AI and LLM-based solutions.
Requirements
  • Bachelor’s or Master’s degree in Computer Science, Software Engineering or related field.
  • Extensive experience designing and implementing production-grade platforms using Azure Databricks.
  • Expertise in PySpark, including advanced optimization, data skew mitigation and query tuning.
  • Strong programming skills in Python with knowledge of modern software design principles.
  • Practical experience with structured streaming, Delta Lake and Delta Live Tables.
  • Proven experience in Lakehouse migration and modernization using open table formats such as Delta Lake or Apache Iceberg.
  • Proficiency with cloud-native services on Azure and knowledge of multi-cloud environments (AWS or GCP).
  • Hands-on experience with CI/CD and Infrastructure as Code tools (Terraform, GitHub Actions, Jenkins).
  • Strong leadership ability to guide teams, define epics/user stories and ensure delivery in agile environments.
  • Excellent communication and stakeholder management skills for both technical and non-technical audiences.
Nice to have
  • Experience operationalizing LLM or Generative AI workflows in Databricks pipelines.
  • Familiarity with frameworks like LangChain, LlamaIndex or Databricks ML/MosaicML.
  • Knowledge of AI governance, security practices and enterprise integration controls.
  • Background in financial trading data or related domains.
  • Official Databricks certifications such as Certified Data Engineer Professional or Apache Spark Developer.
We offer
  • EPAM Employee Stock Purchase Plan (ESPP).
  • Protection benefits including life assurance, income protection and critical illness cover.
  • Private medical insurance and dental care.
  • Employee Assistance Program.
  • Competitive group pension plan.
  • Cyclescheme, Techscheme and season ticket loans.
  • Various perks such as free Wednesday lunch in-office, on-site massages and regular social events.
  • Learning and development opportunities including in-house training and coaching, professional certifications, and courses.
  • If otherwise eligible, participation in the discretionary annual bonus program.
  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program.

Data Engineer - Tech Lead (Databricks, Pyspark) employer: EPAM Systems, Inc.

EPAM Systems, Inc. is an exceptional employer that fosters a dynamic and inclusive work culture, offering employees the flexibility of a hybrid working model in the vibrant city of London. With a strong focus on professional development, team members are encouraged to grow their skills and advance their careers while contributing to impactful digital transformation projects in the energy sector. Joining EPAM means being part of a forward-thinking company that values innovation and collaboration, making it an ideal place for those seeking meaningful and rewarding employment.

EPAM Systems, Inc.

Contact Details:

EPAM Systems, Inc. Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Data Engineer - Tech Lead (Databricks, Pyspark)

Get Involved in Data Science Meetups

Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like EPAM Systems, Inc.!

Show Off Your Projects

Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Data Engineer - Tech Lead (Databricks, Pyspark) at EPAM Systems, Inc..

Leverage Professional Networks

Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like EPAM Systems, Inc..

Apply Directly through Our Website

When you find a suitable opening like Data Engineer - Tech Lead (Databricks, Pyspark) at EPAM Systems, Inc., make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!

We think you need these skills to ace Data Engineer - Tech Lead (Databricks, Pyspark)

SQL
Problem-Solving Skills
Python
Data Pipeline Development
Data Engineering
Communication Skills
API Integration

Some tips for your application 🫡

Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!

Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!

Craft a Tailored Cover Letter:For a full-time role at EPAM Systems, Inc., your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.

Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at EPAM Systems, Inc.. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!

How to prepare for a job interview at EPAM Systems, Inc.

Brush Up on Your Statistics

For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!

Showcase Your Projects

Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!

Get Comfortable with Python and R

Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at EPAM Systems, Inc.!

Prepare for Case Studies

Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.