AI Inference Engineer β€” GPU Performance & Scale

AI Inference Engineer β€” GPU Performance & Scale

Full-Time No working from home possible
Perplexity AI

Perplexity AI seeks an AI Inference Engineer to build and optimize the inference engine behind every Perplexity query. You will support dozens of model architectures at scale with tight latency and cost budgets.

The role involves porting CUDA kernels to CuTe DSL, developing a Rust-based serving runtime, and enhancing performance across the stack from request scheduling to KV-cache management. You will work on GPU programming and ML inference optimizations within a fast-moving team.

#J-18808-Ljbffr

AI Inference Engineer β€” GPU Performance & Scale employer: Perplexity AI

Perplexity is an exceptional employer, offering a dynamic work environment in the heart of London where innovation thrives. With a strong focus on employee growth and collaboration, team members are encouraged to push the boundaries of search technology while enjoying competitive compensation and equity options. The company's rapid expansion and backing from leading investors create unique opportunities for meaningful contributions and career advancement.

Perplexity AI

Contact Details:

Perplexity AI Recruitment Team