At a Glance
- Tasks: Build and scale data pipelines for cutting-edge video generation models.
- Company: Join Cantina Labs, a pioneering social AI company transforming storytelling.
- Benefits: Enjoy competitive salary, generous equity, and extensive paid time off.
- Other info: Dynamic team environment with opportunities for growth and impact.
- Why this job: Be at the forefront of AI innovation and shape the future of creativity.
- Qualifications: 3+ years in ML or data engineering, strong Python skills required.
The predicted salary is between 170000 - 225000 £ per year.
About Cantina
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism.
We bring characters to life, transforming how people tell stories, connect, and create.
We build and power ecosystems.
Cantina, our flagship social AI platform, is just the beginning.
About Cantina
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism.
We bring characters to life, transforming how people tell stories, connect, and create.
We build and power ecosystems.
Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
About The Role
We are looking for a new Member of Technical Staff to build and scale the data pipelines behind our large video generation models.
This role is focused on collecting large amounts of relevant video data, preparing high-quality training samples, and developing robust preprocessing, filtering, and parsing workflows.
You'll orchestrate annotation pipelines across platforms such as MTurk and own the full lifecycle of training data, from raw ingestion to clean, model-ready samples that directly drive quality improvements.
This role sits at the intersection of data engineering and ML research, making it central to how we turn messy real-world data into the fuel that moves our models forward.
- What You’ll Do
- Build and maintain data pipelines for large video generation models, including data ingestion, parsing, filtering, preprocessing, and dataset curation at scale, using tools such as AWS S3 and Dynamo DB.
- Design and run annotation workflows across platforms such as MTurk, Prolific, and Mechanical Turk, including task design, quality control, and label validation.
- Train, evaluate, and improve smaller supporting models used for data filtering, quality assessment, preprocessing, or other parts of the ML pipeline.
- Partner closely with research and engineering teams to turn experimental workflows into scalable, repeatable systems that support model training and evaluation.
- Own data quality across the pipeline by identifying bottlenecks, failure modes, and low-quality sources, and continuously improving tooling and processes.
- Build internal tools and automation that make it easier to prepare datasets, launch annotation jobs, monitor outputs, and support model development end to end.
- Drive larger pipeline projects from start to finish, such as new dataset creation efforts or upgrades to labeling and preprocessing infrastructure.
- Work within a Kubernetes-based training infrastructure, ensuring datasets are properly prepared, formatted, and delivered to training clusters.
- Profile and optimize research model inference scripts used in preprocessing steps, ensuring that model-driven filtering and transformation stages run within practical time and cost constraints when applied to large-scale raw data.
- What You’ll Bring
- 3+ years of experience in machine learning, applied ML, data pipelines, or related engineering roles, ideally working on large-scale multimodal, video, or vision-based systems.
- Strong programming skills in Python and solid experience building reliable data processing and preprocessing pipelines for ML workflows.
- Hands-on experience preparing training data for ML models, including parsing, filtering, dataset curation, quality control, and large-scale data handling using tools such as AWS S3 and Dynamo DB.
- Familiarity with annotation and labeling workflows, including task design, vendor or crowd-platform orchestration such as MTurk or Prolific, and methods for ensuring label quality.
- Experience working with Kubernetes for orchestrating distributed workloads, including data preprocessing, pipeline execution, and dataset delivery to training clusters.
- Comfort working across cloud and on-demand compute environments such as AWS and Run Pod, with the ability to port and optimize pipelines across infrastructure.
- Familiarity with distributed data processing frameworks and experience designing systems that operate reliably at scale across many nodes or workers.
- Working knowledge of Py Torch and the broader deep learning stack, with the ability to read, debug, and optimize research model inference code for use in production preprocessing pipelines.
- Ability to work cross-functionally with research and engineering teams and translate experimental ideas into robust, scalable systems.
- Bachelor's, Master's, or Ph D in Computer Science, Machine Learning, Engineering, Mathematics, or a related technical field; experience in generative video, computer vision, or multimodal ML is strongly preferred.
- Bonus: Experience training, evaluating, or fine-tuning smaller ML models used for classification, filtering, ranking, quality assessment, or other supporting tasks in an ML pipeline.
Compensation
The anticipated annual base salary range for this role is between $200,000-$260,000 (€170,000-€225,000).
When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
Benefits For U. S.-based Roles
- Competitive salary and generous company equity
- Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
• 42 days of paid time off, including
- 15 PTO days
- 10 sick days
- 15 company holidays
- 2 floating holidays
- Generous parental leave & fertility support
- 401(k) retirement savings plan
- Lifestyle spending account – $500/month to use however you’d like
- Complimentary lunch and snacks for in-office employees
• One Medical membership, and more!
#J-18808-Ljbffr
Member of Technical Staff, Data & ML Infrastructure for Video Models employer: Cantina Labs
Cantina Labs is an exceptional employer that fosters a vibrant work culture focused on innovation and creativity in the realm of social AI. With generous benefits including competitive salaries, extensive paid time off, and comprehensive health coverage, employees are empowered to thrive both personally and professionally. The collaborative environment encourages growth and development, making it an ideal place for those passionate about shaping the future of AI and storytelling.
StudySmarter Expert Advice🤫
We think this is how you could land Member of Technical Staff, Data & ML Infrastructure for Video Models
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Cantina Labs!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Member of Technical Staff, Data & ML Infrastructure for Video Models at Cantina Labs.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Cantina Labs.
✨Apply Directly through Our Website
When you find a suitable opening like Member of Technical Staff, Data & ML Infrastructure for Video Models at Cantina Labs, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Member of Technical Staff, Data & ML Infrastructure for Video Models
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at Cantina Labs, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Cantina Labs. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at Cantina Labs
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Cantina Labs!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.