At a Glance
- Tasks: Drive innovative multimodal generative models for AI video communication.
- Company: Join Synthesia, a leading AI video platform valued at $4 billion.
- Benefits: Competitive salary, remote work options, and opportunities for professional growth.
- Other info: Collaborative environment with a focus on cutting-edge technology and research.
- Why this job: Shape the future of interactive AI and make a real impact in visual communication.
- Qualifications: Experience in generative modelling and strong skills in PyTorch required.
The predicted salary is between 80000 - 100000 £ per year.
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations.
About the role
Synthesia's long-term vision is to build the best human-interactive models: systems that don't just talk to people, but perceive and respond to them, reacting to a user's actions and emotions, not only their words. Today over 60,000 businesses rely on our platform, and the next leap in what we can offer them depends on models that combine text, audio, and video into a single real-time interactive experience. As a Staff Research Engineer, you'll join the Voice team within our 40+ person R&D department, but your scope will extend well beyond voice. You'll help define and drive that broader vision across teams, proposing ambitious research directions and taking direct ownership of the design and implementation of its most critical components. You'll work directly with our voice lead and collaborate tightly with our video teams and other senior members of the org.
What you'll do
- Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons.
- Propose novel multi-modal system architectures (especially text and voice).
- Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
- Design solutions that reinforce emotional expressiveness and natural interaction.
- Implement and bring designs to life, from pretraining through post-training.
- Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness.
- Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
- Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
- Curate new datasets to complement existing data.
- Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality.
- Ship models to production with optimised runtime to serve customers, and address their feedback thereafter.
You’ll thrive in this role if you have
- The ability to bring novel ideas and designs that advance the field of interactive multimodal systems.
- Strong understanding of generative modelling, ideally applied to sequential or multimodal data.
- Hands-on experience with large language models or similar transformer-based architectures.
- High proficiency in PyTorch, including distributed training and model optimisation.
- A solid grasp of time-series modelling and tokenisation, preferably in the context of audio, speech, or video.
- A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently.
- Proven experience training deep learning models end-to-end, from data preparation through evaluation.
- Strong general software engineering skills, enabling contributions to a large, shared research infrastructure.
Particularly relevant experience
- Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it.
- Working on conversational or interactive systems where latency, responsiveness, and user experience were first-class constraints, not afterthoughts.
- Working on LLMs with large scale trainings leading to models with decent reasoning capabilities.
- Owning a research problem end to end: from architecture proposal through pretraining, post-training, and production deployment.
- Collaborating across modalities or teams (e.g. audio and video, or research and product) to ship a unified system.
Bonus points for
- Experience with real-time or streaming architectures.
- Familiarity with state-of-the-art architectures in audio and speech generation, such as diffusion models, neural codecs, flow-matching models, or autoregressive decoders.
- Excellence in one or more of the following modalities: voice, text, video.
- Evidence of original research contributions, such as publications or open-source work at top-tier venues (e.g. NeurIPS, CVPR, ICML, ICLR, Interspeech).
Staff Research Engineer - Multimodal Generative Modelling employer: SLAMcore
At SLAMcore, we pride ourselves on being an exceptional employer that values innovation and collaboration. Our supportive work culture fosters personal and professional growth, offering ample opportunities for career advancement while ensuring a meaningful impact on our users' experiences. Located in a vibrant tech hub, we provide a dynamic environment where your contributions are recognised and celebrated, making every day rewarding.
StudySmarter Expert Advice🤫
We think this is how you could land Staff Research Engineer - Multimodal Generative Modelling
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like SLAMcore!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Staff Research Engineer - Multimodal Generative Modelling at SLAMcore.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like SLAMcore.
✨Apply Directly through Our Website
When you find a suitable opening like Staff Research Engineer - Multimodal Generative Modelling at SLAMcore, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Staff Research Engineer - Multimodal Generative Modelling
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at SLAMcore, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at SLAMcore. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at SLAMcore
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at SLAMcore!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.