Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.
You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.
#J-18808-Ljbffr
AI Inference Engineer (Staff) - GPU ML Systems + Equity employer: Perplexity
Perplexity is an exceptional employer that fosters a collaborative and innovative work culture, where your contributions directly impact the evolution of cutting-edge search technologies. With a hybrid work model in vibrant cities like Belgrade, London, or Berlin, employees enjoy flexibility alongside opportunities for professional growth and development in a dynamic environment. Join us to be part of a team that values creativity and excellence, while also providing competitive benefits and a supportive atmosphere for personal and career advancement.