At a Glance
- Tasks: Architect and manage data pipelines for high-resolution medical imagery.
- Company: TileBio, a pioneering company in pathology AI.
- Benefits: Competitive salary, equity options, and impactful work culture.
- Other info: Dynamic environment with opportunities for innovation and growth.
- Why this job: Join us to revolutionise biological discovery with cutting-edge data engineering.
- Qualifications: Strong Python skills and experience with large datasets required.
The predicted salary is between 35000 - 60000 £ per year.
TileBio is building a foundation model of pathology. We treat histology as a language of tissue and train models directly on raw, unlabelled whole slide images (WSIs) to learn the underlying structure of biology. Our system converts tissue into a sequence of learned tokens and trains transformer models to understand disease at scale. Data is our lifeblood, and we are looking for a Data Engineer to build the refinery. At TileBio, "Data Engineering" isn't just about moving rows in a database; it’s about architecting the flow of terabytes of high-resolution medical imagery. You will build and maintain the infrastructure that fuels our foundation model training. You will own the lifecycle of a Whole Slide Image from the moment it leaves a scanner or a partner’s server to the moment it becomes a versioned, reproducible training set for our AI team.
What You'll Do
- Architect the ingestion and preprocessing of massive histopathology datasets (SVS, NDPI, iSyntax)
- Build and refine automated, high-throughput pipelines for tiling, filtering, and metadata extraction
- Implement rigorous dataset versioning so every training experiment is traceable and repeatable
- Coordinate data movement between local high-performance storage, cluster environments (SLURM), and cloud buckets
- Monitor and resolve quality issues, ensuring our models learn from the highest quality biological signals
- Work closely with Deep Learning Engineers to optimise data loading and throughput for multi-GPU training
What We're Looking For
- Strong Python skills with the ability to write clean, performant, and maintainable code
- Hands-on experience managing large scientific, imaging, or unstructured datasets
- Experience with storage architectures such as ZFS, object storage, or cloud-native solutions
- Experience building and deploying automated data pipelines (e.g., Prefect, Dagster, Airflow, or custom-built)
- Deep understanding of data versioning (e.g., DVC, LakeFS) and why it matters for ML
Nice to Have
- Experience with WSI formats or medical imaging libraries (OpenSlide, CuCIM)
- Experience with SLURM or other cluster-based scheduling systems
- Familiarity with Azure or AWS environments for hybrid-cloud workflows
- Strong systems thinking with the ability to design for "the next 100TB" today
What We Offer
- Salary of £35,000 to £60,000 (depending on experience) and equity eligibility
- Opportunity to build the core infrastructure that enables biological discovery at a global scale
- A culture that optimises for impact, takes ownership, and delivers
- An environment that challenges ideas, not people, and values speed, clarity, and rigour
AI/Data Engineer – Data Infrastructure in Glasgow employer: TileBio
TileBio is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the cutting-edge field of AI and pathology. Located in Glasgow, employees benefit from competitive salaries, equity options, and ample opportunities for professional growth, all while contributing to meaningful advancements in healthcare technology.
StudySmarter Expert Advice🤫
We think this is how you could land AI/Data Engineer – Data Infrastructure in Glasgow
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like TileBio!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like AI/Data Engineer – Data Infrastructure at TileBio.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like TileBio.
✨Apply Directly through Our Website
When you find a suitable opening like AI/Data Engineer – Data Infrastructure at TileBio, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace AI/Data Engineer – Data Infrastructure in Glasgow
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at TileBio, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at TileBio. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at TileBio
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at TileBio!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.