AI Inference Engineer (Staff) - GPU ML Systems + Equity in London

AI Inference Engineer (Staff) - GPU ML Systems + Equity in London

London Full-Time On-site
P

Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.

You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.

#J-18808-Ljbffr

AI Inference Engineer (Staff) - GPU ML Systems + Equity in London employer: Perplexity

Perplexity is an exceptional employer that fosters a collaborative and innovative work culture, where your contributions directly impact the evolution of cutting-edge search technologies. With a hybrid work model in vibrant cities like Belgrade, London, or Berlin, employees enjoy flexibility alongside opportunities for professional growth and development in a dynamic environment. Join us to be part of a team that values creativity and excellence, while also providing competitive benefits and a supportive atmosphere for personal and career advancement.

P

Contact Details:

Perplexity Recruitment Team