Principal Data Architect in Glasgow

Principal Data Architect in Glasgow

Glasgow Full-Time 80000 - 100000 £ / year (est.) Home office (partial)
C

At a Glance

  • Tasks: Design and implement data architecture for cutting-edge AI-driven chemistry projects.
  • Company: Join Chemify, a pioneering tech company transforming the future of chemistry.
  • Benefits: Competitive pay, flexible work options, and the chance to shape groundbreaking technology.
  • Other info: Be part of a rapidly scaling team with significant career growth opportunities.
  • Why this job: Make a real impact in science and medicine while working with innovative technologies.
  • Qualifications: Extensive experience in data architecture, Python, SQL, and graph databases.

The predicted salary is between 80000 - 100000 £ per year.

Location: Glasgow or Remote

Workstyle: Hybrid

Reports to: CTO

About Chemify:

Chemify is revolutionising chemistry. We are creating a future where the synthesis of previously unimaginable molecules, drugs, and materials is instantly accessible. By combining AI, robotics, and the world’s largest continually expanding database of chemical programs, we are accelerating chemical discovery to improve quality of life and extend the reach of humanity.

The Role:

At Chemify, robots run real chemical experiments around the clock, and every spectrum, video frame, and sensor reading is gathered to decide what to synthesise next. We're bringing in a Principal Data Architect on a focused six-month engagement to set the architecture that turns a firehose of scientific literature and robotic lab data into a clean, governed, ML-ready ecosystem that our scientists and AI models can build on. You'll define the blueprint and stand up the foundations during the engagement, working closely with our ML and data engineering teams so they can carry it forward. We're looking for someone senior enough to make the right calls quickly and pragmatic enough to deliver real systems in the time allocated.

Scope of the Engagement:

Over six months you'll own how data flows across Chemify, including how it's stored, synchronised, governed, and shared, within real regulatory and contractual limits. Concretely, you'll aim to:

  • Set the AI-ready foundation. Partner with our ML and data engineering teams to design ingestion and orchestration that captures scientific and operational data in a form models can use immediately. You'll define and stand up the initial data Lakehouse on AWS for the raw data coming off our robots, and design a versioned feature store that standardises chemical descriptors so there's a tight loop between lab execution and the discovery AI.
  • Model chemistry as data. Architect graph models for complex reaction networks, define the semantic ontologies that let AI reason across different chemical data types, and design vector search (e.g. pgvector, Pinecone) for similarity lookups across large molecular sets.
  • Design the telemetry pipeline. Specify streaming ingestion for high-frequency robot and sensor data, with zero loss of the "negative data" from failed experiments that's so valuable for training, plus a distributed pattern that keeps local lab "edge" data in sync with central training clusters without sacrificing consistency.
  • Set the governance guardrails. Define the tenancy and partitioning models that guarantee strict IP isolation for enterprise clients, secure sharing patterns for research partners, and the architectural groundwork for SOC 2 and ISO 27001 readiness.

We'll agree concrete deliverables and milestones with you up front, so success is clearly defined on both sides.

What you will bring:

This is a senior contract with little onboarding runway, so we're looking for someone who's done most of this before:

  • Extensive experience as a data architect, with production-grade Python and SQL
  • Deep PostgreSQL experience, ideally on AWS RDS
  • Hands-on work with graph databases and modelling highly relational domains
  • A proven track record designing distributed data systems across multiple teams, services, or locations
  • Demonstrable experience with data governance, data contracts, tenancy and segregation, replication and consistency patterns, and secure cross-organisation data sharing
  • The ability to set direction independently, communicate clearly with both engineers and scientists, and deliver within a fixed timeframe
  • Familiarity with modern data stacks: lakehouses, streaming, batch and real-time pipelines

Desirable:

  • High-throughput telemetry, IoT, or industrial time-series systems
  • Prior involvement in SOC 2 or ISO 27001 programmes
  • Exposure to scientific, chemical, or manufacturing data
  • Chemistry or AI drug-discovery knowledge

Why Join Chemify?

Impact: You will help build the infrastructure that enables digital chemistry at scale — accelerating discovery, improving reproducibility, and unlocking new possibilities in science and medicine.

Autonomy: Reporting directly to the CTO, you will have meaningful influence over the technical direction and data strategy of a Series B deep-tech rocket ship.

Ambition: We are scaling rapidly, investing in world-class infrastructure, and tackling problems that sit at the frontier of robotics, AI, and chemistry. You will have the resources and mandate to build the right foundations for the future.

Principal Data Architect in Glasgow employer: Chemify

Chemify is an exceptional employer, offering a unique opportunity for Senior Electronics Engineers to work at the forefront of chemistry and robotics in Glasgow. With a strong emphasis on autonomy and impact, you will have the chance to design electronics that directly influence groundbreaking advancements in chemical synthesis. The collaborative work culture fosters innovation and growth, ensuring that your contributions are not only valued but also pivotal in shaping the future of technology.

C

Contact Details:

Chemify Recruitment Team

We think you need these skills to ace Principal Data Architect in Glasgow

Data Architecture
Python
SQL
PostgreSQL
Graph Databases
Distributed Data Systems
Data Governance