At a Glance
- Tasks: Drive innovative multimodal generative modelling for AI video interactions.
- Company: Join Synthesia, the leading AI video platform trusted by Fortune 100 companies.
- Benefits: Competitive salary, flexible working, and opportunities for professional growth.
- Other info: Collaborative environment with a focus on impactful research and development.
- Why this job: Be at the forefront of AI technology, shaping the future of visual communication.
- Qualifications: Experience in generative modelling and proficiency in PyTorch required.
The predicted salary is between 80000 - 100000 £ per year.
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations.
Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow.
About the role: Synthesia's long-term vision is to build the best human-interactive models: systems that don't just talk to people, but perceive and respond to them, reacting to a user's actions and emotions, not only their words. Today over 60,000 businesses rely on our platform, and the next leap in what we can offer them depends on models that combine text, audio, and video into a single real-time interactive experience.
As a Staff Research Engineer, you'll join the Voice team within our 40+ person R&D department, but your scope will extend well beyond voice. You'll help define and drive that broader vision across teams, proposing ambitious research directions and taking direct ownership of the design and implementation of its most critical components. You'll work directly with our voice lead and collaborate tightly with our video teams and other senior members of the org.
Concretely, you'll work on voice to voice models that produce text and voice simultaneously. Models that can reason, interrupt the user with back channeling and talk. Models that feel like you are having a natural conversation with, without the feeling of turn taking that is dominant in current speech to speech models. Your role would be to partner in defining a roadmap, implementing it and shipping the outcomes to product.
What you'll do:
- Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons.
- Propose novel multi-modal system architectures (especially text and voice).
- Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
- Design solutions that reinforce emotional expressiveness and natural interaction.
- Implement and bring designs to life, from pretraining through post-training.
- Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness.
- Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
- Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
- C curate new datasets to complement existing data.
- Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality.
- Ship models to production with optimised runtime to serve customers, and address their feedback thereafter.
You’ll thrive in this role if you have:
- The ability to bring novel ideas and designs that advance the field of interactive multimodal systems.
- Strong understanding of generative modelling, ideally applied to sequential or multimodal data.
- Hands-on experience with large language models or similar transformer-based architectures.
- High proficiency in PyTorch, including distributed training and model optimization.
- A solid grasp of time-series modeling and tokenization, preferably in the context of audio, speech, or video.
- A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently.
- Proven experience training deep learning models end-to-end, from data preparation through evaluation.
- Strong general software engineering skills, enabling contributions to a large, shared research infrastructure.
Particularly relevant experience:
- Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it.
- Working on conversational or interactive systems where latency, responsiveness, and user experience were first-class constraints, not afterthoughts.
- Working on LLMs with large scale trainings leading to models with decent reasoning capabilities.
- Owning a research problem end to end: from architecture proposal through pretraining, post-training, and production deployment.
- Collaborating across modalities or teams (e.g. audio and video, or research and product) to ship a unified system.
Bonus points for:
- Experience with real-time or streaming architectures.
- Familiarity with state-of-the-art architectures in audio and speech generation, such as diffusion models, neural codecs, flow-matching models, or autoregressive decoders.
- Excellence in one or more of the following modalities: voice, text, video.
- Evidence of original research contributions, such as publications or open-source work at top-tier venues (e.g. NeurIPS, CVPR, ICML, ICLR, Interspeech).
Staff Research Engineer - Multimodal Generative Modelling in London employer: Synthesia
At Synthesia, we pride ourselves on being an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the heart of Greater London. Our commitment to employee growth is evident through continuous learning opportunities and the chance to work on cutting-edge AI technologies, making every day at Synthesia both meaningful and rewarding.
StudySmarter Expert Advice🤫
We think this is how you could land Staff Research Engineer - Multimodal Generative Modelling in London
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Synthesia!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Staff Research Engineer - Multimodal Generative Modelling at Synthesia.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Synthesia.
✨Apply Directly through Our Website
When you find a suitable opening like Staff Research Engineer - Multimodal Generative Modelling at Synthesia, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Staff Research Engineer - Multimodal Generative Modelling in London
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at Synthesia, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Synthesia. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at Synthesia
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Synthesia!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.