Senior ML Research Engineer, Virtual Cell

Senior ML Research Engineer, Virtual Cell

Full-Time 59400 - 72600 £ / year (est.) Home office (partial)
Sandboxaq

At a Glance

  • Tasks: Develop cutting-edge ML models to revolutionise drug discovery and impact healthcare.
  • Company: Join a high-growth AI company tackling global challenges with a collaborative culture.
  • Benefits: Enjoy competitive pay, comprehensive benefits, and flexible work arrangements.
  • Other info: Be part of a diverse team committed to innovation and continuous growth.
  • Why this job: Make a real difference in life sciences while advancing your career in a dynamic environment.
  • Qualifications: Experience in machine learning and large-scale data management; a scientific background is preferred.

The predicted salary is between 59400 - 72600 £ per year.

About SandboxAQ

SandboxAQ is a high-growth company delivering AI solutions that address some of the world's greatest challenges. The company’s Large Quantitative Models (LQMs) power advances in life sciences, financial services, navigation, cybersecurity, and other sectors. We are a global team that is tech-focused and includes experts in AI, chemistry, cybersecurity, physics, mathematics, medicine, engineering, and other specialties. The company emerged from Alphabet Inc. as an independent, growth capital-backed company in 2022, funded by leading investors and supported by a braintrust of industry leaders.

At SandboxAQ, we’ve cultivated an environment that encourages creativity, collaboration, and impact. By investing deeply in our people, we’re building a thriving, global workforce poised to tackle the world's epic challenges. Join us to advance your career in pursuit of an inspiring mission, in a community of like-minded people who value entrepreneurialism, ownership, and transformative impact.

The Opportunity

The AI Sim R&D team builds leading-edge ML and physics-based models ("LQMs") to advance drug discovery. Within this team, AQCell is our virtual cell platform: it takes a cell representation (e.g. basal gene expression) and a perturbation descriptor (e.g. SMILES, dose, and time), predicts the resulting transcriptomic response, and maps that response through pathway activity to functional endpoints such as cell viability, IC50, and toxicity dose-response — helping drug discovery scientists understand the biological repercussions of a compound across the cell, not just whether it binds its target.

As a Machine Learning Engineer on AQCell, you will build and maintain the models and data infrastructure that power this pipeline. You will work across large, heterogeneous transcriptomic and functional-endpoint datasets (e.g. LINCS L1000, GDSC, Tahoe-100M, and DILImap), train and evaluate expression-perturbation and cell-viability prediction models over them, and help harden our evaluation pipeline and baselines so we can trust and improve model performance over time. This role sits at the intersection of machine learning and biology, and offers the opportunity to shape how virtual cell models are trained, validated, and scaled toward real drug discovery decisions.

Key Responsibilities

  • Model Development: Build, train, and maintain machine learning models for expression-perturbation prediction (e.g. transcriptomic response to a drug perturbation) and downstream functional-endpoint prediction (e.g. cell viability, IC50, toxicity dose-response).
  • Large-Scale Dataset Management: Acquire, harmonize, and manage large-scale biological datasets (e.g. LINCS L1000, GDSC, Tahoe-100M, DILImap) — including schema harmonization, normalization, and de-duplication across cell, drug, and assay identifiers — and manage model training pipelines over these pooled datasets.
  • Evaluation & Baselines: Contribute to automating and hardening the end-to-end evaluation pipeline, including implementing robust statistical baselines (e.g. cell- and drug-conditioned mean baselines) to rigorously benchmark model performance.
  • Research Translation: Translate ideas from the scientific literature (e.g. transformer-based perturbation models, knowledge graph and GNN-based embeddings) into working, well-tested code integrated into our modeling framework.
  • Cross-Functional Collaboration: Partner with computational biologists, software engineers, and product stakeholders to ensure models are grounded in sound biology and are usable in real drug discovery workflows.
  • Communication: Clearly document methods, assumptions, and results, and communicate findings to both technical and non-technical stakeholders.

Essential Skills & Experience

  • Academic Foundation: Bachelor's degree in a scientific or quantitative field (Computer Science, Physics, Mathematics, Biology, Chemistry, or related); an advanced degree (MS or PhD) is preferred.
  • Applied ML in Science: Demonstrated experience building and maintaining machine learning models in a scientific discipline in an industry setting, including taking models from prototype through validation and maintenance; experience in a life sciences setting is preferred.
  • Large-Scale Data Management: Required experience managing large-scale datasets and managing the training of models over them, including data ingestion, cleaning, versioning, and pipeline maintenance at scale.
  • Software Engineering: Strong Python programming skills and experience with modern ML frameworks (e.g. PyTorch, JAX) and experiment tracking/data versioning tools (e.g. Weights & Biases).
  • Scientific Rigor: Ability to design sound evaluation methodology (e.g. train/test splitting strategies, held-out generalization tests) and to critically interpret model performance against meaningful baselines.

Highly Desired Skills & Experience

  • Bioinformatics & Computational Biology: Experience with bioinformatics and computational biology data analysis tools, particularly transcriptomics harmonization and normalization tools (e.g. batch correction, pseudobulking, gene ID mapping/standardization).
  • Domain Datasets: Familiarity with public perturbation or drug-sensitivity datasets such as LINCS L1000, GDSC, DepMap, or single-cell perturbation atlases.
  • Knowledge Graphs & GNNs: Experience with knowledge graph embeddings or graph neural networks applied to drugs, targets, or cells.
  • Chemoinformatics: Familiarity with cheminformatics representations (SMILES, InChI keys, PubChem/Cellosaurus identifiers) used to harmonize drug and cell metadata across datasets.
  • Collaborative Innovation: Experience working in interdisciplinary environments where AI intersects with the biological sciences.

Why Join Us?

  • We offer competitive compensation, a comprehensive benefits package, and opportunities for professional growth.
  • Compensation: Competitive base salary, performance-based incentives or bonuses (where applicable), and equity participation.
  • Benefits: Comprehensive medical, dental, and vision coverage for employees and dependents with generous employer premium contributions, retirement savings with company matching, paid parental leave, and inclusive family-building benefits.
  • Work-Life Balance: Flexible paid time off, company-wide seasonal breaks, and support for flexible work arrangements that enable sustainable performance.
  • Career Development: Opportunities for continuous learning and growth through on-the-job development, cross-functional collaboration, and access to internal learning and development programs.

SandboxAQ Welcomes All

We are committed to fostering a culture of belonging and respect, where diverse perspectives are actively sought and valued. Our multidisciplinary environment provides ample opportunity for continuous growth - working alongside humble, empowered, and ambitious colleagues ready to tackle epic challenges.

Equal Employment Opportunity: All qualified applicants will receive consideration regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, or Veteran status.

Accommodations: We provide reasonable accommodations for individuals with disabilities in job application procedures for open roles. If you need such an accommodation, please let a member of our Recruiting team know.

Read: Guidance for candidates on using AI Tools in interviews

Senior ML Research Engineer, Virtual Cell employer: Sandboxaq

At SandboxAQ, we pride ourselves on being an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the heart of the UK. Our commitment to employee growth is evident through our investment in cutting-edge technology and continuous learning opportunities, ensuring that you can thrive as a Senior ML Research Engineer while contributing to groundbreaking advancements in life sciences AI. Join us to be part of a team that values creativity and impact, where your work directly influences the future of drug and materials discovery.

Sandboxaq

Contact Details:

Sandboxaq Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Senior ML Research Engineer, Virtual Cell

Get Involved in Data Science Meetups

Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Sandboxaq!

Show Off Your Projects

Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Senior ML Research Engineer, Virtual Cell at Sandboxaq.

Leverage Professional Networks

Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Sandboxaq.

Apply Directly through Our Website

When you find a suitable opening like Senior ML Research Engineer, Virtual Cell at Sandboxaq, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!

We think you need these skills to ace Senior ML Research Engineer, Virtual Cell

Machine Learning Model Development
Large-Scale Dataset Management
Statistical Evaluation Methodology
Python Programming
PyTorch
JAX
Data Ingestion and Cleaning

Some tips for your application 🫡

Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!

Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!

Craft a Tailored Cover Letter:For a full-time role at Sandboxaq, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.

Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Sandboxaq. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!

How to prepare for a job interview at Sandboxaq

Brush Up on Your Statistics

For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!

Showcase Your Projects

Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!

Get Comfortable with Python and R

Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Sandboxaq!

Prepare for Case Studies

Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.