At a Glance
- Tasks: Research and develop cutting-edge speech synthesis algorithms and systems.
- Company: Join Tencent's Lightspeed Studios, a leader in innovative game development.
- Benefits: Competitive salary, diverse work culture, and opportunities for professional growth.
- Other info: Collaborative environment with a focus on innovation and diversity.
- Why this job: Make an impact in the gaming industry with advanced technology and creative storytelling.
- Qualifications: Ph.D. in relevant fields and strong skills in speech/audio processing.
The predicted salary is between 63000 - 77000 Β£ per year.
Business Unit
LIGHTSPEED STUDIOS is made up of passionate players who advance the art & science of game development through great stories, great gameplay, and advanced technology.
We are focused on bringing next generation experiences to gamers who want to enjoy them anywhere, anytime, across multiple genres and devices.
Business Unit
LIGHTSPEED STUDIOS is made up of passionate players who advance the art & science of game development through great stories, great gameplay, and advanced technology.
We are focused on bringing next generation experiences to gamers who want to enjoy them anywhere, anytime, across multiple genres and devices.
About The Hiring Team
Lightspeed Tech Center is a R&D department under Lightspeed Studios which develop PUBG Mobile and other high-quality games.
Our Tech Center leads the research, exploration, and discovery of innovative technologies and provides technical services for all games during all phases of life cycle, including engine, audio, QA, AI, next generation game, technical cooperation, etc.
- What The Role Entails
- Research and develop advanced speech synthesis and generation algorithms (e. g., TTS, voice conversion, sound/music generation) based on LLMs, multimodal/omnimodal LLMs.
- Develop and optimize speech and audio synthesis systems for online applications, improving effectiveness, efficiency, and scalability.
- Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction.
- Collaborate cross-functionally with research and engineering teams from prototyping to production.
- Who We Look For
- Ph. D. in Computer Science, Electrical Engineering, Signal Processing, or a closely related field.
- Strong foundation in speech/audio processing and modern generative models (e. g., diffusion, flow matching, autoregressive, codec-based approaches).
- Hands-on experience extending LLMs to speech/audio modalities (e. g., speech tokenizers, multimodal adapters, speech-text joint training).
Experience with full-duplex or streaming spoken dialogue systems and real-time interaction modeling is a plus.
- Proficient in Python and deep learning frameworks (e. g., Py Torch); experience with distributed training is a plus.
- Track record of publications at top-tier venues (e. g., ICML, Neur IPS, ICLR, ACL, ICASSP, Interspeech).
- Equal Employment Opportunity at Tencent
As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community.
We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.
#J-18808-Ljbffr
Senior Researcher, Speech Synthesis and Multimodal LLM employer: Tencent
Tencent offers a dynamic and innovative work environment where interns can gain hands-on experience in AI development while collaborating with leading experts in the field. With a strong emphasis on employee growth, Tencent provides numerous opportunities for professional development and mentorship, making it an excellent choice for those looking to kickstart their careers in technology and marketing. Located in a vibrant tech hub, the company fosters a culture of creativity and collaboration, ensuring that every team member's contributions are valued and impactful.