Senior Data Engineer (Metadata & Lineage Integration) in City of London

Senior Data Engineer (Metadata & Lineage Integration) in City of London

City of London Full-Time 63000 - 77000 £ / year (est.) No working from home possible
I

At a Glance

  • Tasks: Build and enhance data systems for AI integration in a leading investment firm.
  • Company: Global investment management company based in London, managing over $228 billion in assets.
  • Benefits: Competitive salary, flexible working options, and opportunities for professional growth.
  • Other info: Dynamic environment with a focus on innovation and collaboration.
  • Why this job: Join a cutting-edge team to shape the future of AI in finance.
  • Qualifications: 5+ years in Python data engineering with strong SQL skills required.

The predicted salary is between 63000 - 77000 £ per year.

Our client is a leading global investment management company headquartered in London. It manages over $228 billion in assets and serves institutional investors, pension funds, wealth managers, and other sophisticated clients worldwide. The firm specializes in quantitative investing, alternative investments, systematic trading strategies, and technology-driven asset management. Data science, machine learning, and AI are core components of its investment and research processes.

As part of our collaboration, we will focus on two foundational capabilities required to enable safe and scalable AI adoption across the enterprise: Agentic Security and AI-Ready Data Foundations. We build the data foundations that make AI useful and safe inside regulated financial firms. The value of AI is capped by the data its agents can reach: if an agent cannot find, interpret, trace or be correctly permissioned against data, the capability is useless, or worse, unsafe. Your job is to close that gap.

This is a hands-on senior role for an excellent Python engineer with strong data-engineering skills who is genuinely comfortable building with AI agents. You will design and build the catalogue, semantic, entitlement and analytical layers that turn large on-premise data estates into something agents can use.

Requirements:

  • 5+ years building production data systems in Python, with strong engineering fundamentals (testing, code review, performance) and solid SQL.
  • Experience building crawlers, harvesters or connector frameworks that extract inventories, schemas, field dictionaries and lineage from databases, filesystems, message platforms and API surfaces, in addition to conventional data pipelines.
  • Event-driven integration with Kafka or similar, including secure producer patterns (mTLS or equivalent) and schema-managed topics.
  • Experience with search and document stores that back catalogue platforms (e.g. Elasticsearch, OpenSearch, MongoDB or similar).
  • Working knowledge of lineage capture and modelling, with OpenLineage or similar as a reference, and readiness to work with proprietary in-house event models.
  • Experience applying LLMs to metadata work, such as drafting descriptions and classifications for human review, including quality evaluation of the generated output.
  • Readiness to work inside another team's codebase, complete components that the team has designed, and contribute through its review process.
  • Fluent English for written and spoken communication with client teams.

Will be a plus:

  • Time-series and tick stores (e.g. kdb+ or similar columnar time-series databases), market-data vendor schema APIs, symbology and asset-class concepts (market-data opening).
  • MS SQL Server estates, reporting and BI systems, inventory extraction from application metadata tables (reporting opening).
  • Columnar and lake formats (Parquet or similar), large object stores, orchestration platforms (Airflow or similar).
  • Working-level knowledge of graph databases; data contracts and data quality frameworks; catalogue platforms (DataHub or similar) on the ingestion side.
  • Day-to-day use of AI coding agents; building data services consumed by AI agents.
  • Experience in financial services or other regulated on-premise environments.

Responsibilities:

  • Build extraction, enrichment and registration paths that populate domain catalogues from live estates (databases, time-series stores, streaming platforms, filesystems, internal and vendor APIs) and connect them to the client's central catalogue through its existing mechanisms.
  • Extend an existing scraping capability from bare dataset and symbol inventories to full metadata: descriptions, field-level dictionaries, date ranges, asset-class and cadence tags, vendor provenance.
  • Load vendor schema metadata at scale, through vendor APIs, into a persistent internal knowledge base designed for step-by-step disclosure to humans and LLMs.
  • Seed report and dataset inventories from existing application metadata tables and ETL sources; combine them with LLM-drafted descriptions approved by stewards.
  • Integrate lineage into the client's lineage backend across batch, streaming and cross-system report chains, completing the client's existing registration designs.
  • Implement the federation contract defined by the architect: stable identities, ownership, hierarchy, links and availability state exposed by each local catalogue to the central layer.
  • Keep metadata current through scheduled and event-driven refresh, with explicit staleness detection and quality signals.

Senior Data Engineer (Metadata & Lineage Integration) in City of London employer: Intellias

As a Senior Python Engineer at our client, a leading global investment management company in London, you will thrive in a dynamic and innovative environment that prioritises technology-driven asset management. The firm offers exceptional benefits, a collaborative work culture, and ample opportunities for professional growth, ensuring that employees are equipped to excel in their roles while contributing to cutting-edge AI and data engineering projects. With a focus on secure and scalable AI adoption, this role provides a unique chance to work with large datasets and advanced technologies, making a meaningful impact in the financial services sector.

I

Contact Details:

Intellias Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Senior Data Engineer (Metadata & Lineage Integration) in City of London

Get Involved in Data Science Meetups

Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Intellias!

Show Off Your Projects

Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Senior Data Engineer (Metadata & Lineage Integration) at Intellias.

Leverage Professional Networks

Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Intellias.

Apply Directly through Our Website

When you find a suitable opening like Senior Data Engineer (Metadata & Lineage Integration) at Intellias, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!

We think you need these skills to ace Senior Data Engineer (Metadata & Lineage Integration) in City of London

Python
SQL
Data Engineering
Event-Driven Integration
Kafka
Elasticsearch
OpenSearch

Some tips for your application 🫡

Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!

Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!

Craft a Tailored Cover Letter:For a full-time role at Intellias, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.

Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Intellias. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!

How to prepare for a job interview at Intellias

Brush Up on Your Statistics

For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!

Showcase Your Projects

Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!

Get Comfortable with Python and R

Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Intellias!

Prepare for Case Studies

Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.