Perplexity AI seeks an AI Inference Engineer to build and optimize the inference engine behind every Perplexity query. You will support dozens of model architectures at scale with tight latency and cost budgets.
The role involves porting CUDA kernels to CuTe DSL, developing a Rust-based serving runtime, and enhancing performance across the stack from request scheduling to KV-cache management. You will work on GPU programming and ML inference optimizations within a fast-moving team.
#J-18808-Ljbffr
AI Inference Engineer β GPU Performance & Scale employer: Perplexity AI
Perplexity is an exceptional employer, offering a dynamic work environment in the heart of London where innovation thrives. With a strong focus on employee growth and collaboration, team members are encouraged to push the boundaries of search technology while enjoying competitive compensation and equity options. The company's rapid expansion and backing from leading investors create unique opportunities for meaningful contributions and career advancement.