AI Inference Engineer (Staff) - GPU ML Systems + Equity

AI Inference Engineer (Staff) - GPU ML Systems + Equity

Full-Time No working from home possible
P

Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.

You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.

#J-18808-Ljbffr

AI Inference Engineer (Staff) - GPU ML Systems + Equity employer: Perplexity

Perplexity is an exceptional employer that fosters a collaborative and innovative work culture, where your contributions directly impact the evolution of cutting-edge search technologies. With a hybrid work model in vibrant cities like Belgrade, London, or Berlin, employees enjoy flexibility alongside opportunities for professional growth and development in a dynamic environment. Join us to be part of a team that values creativity and excellence, while also providing competitive benefits and a supportive atmosphere for personal and career advancement.

P

Contact Details:

Perplexity Recruitment Team