AI Inference Engineer (Member of Technical Staff)

AI Inference Engineer (Member of Technical Staff)

Full-Time No working from home possible
Perplexity AI
  • We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL
  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA’s CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents

You understand modern LLM architectures and are able to bring them up reliably in a production environmentDeep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plusYou’ve built and operated production distributed systems under real load - ideally performance-critical onesYou own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on FridayComfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernelsSelf-directed. You do well in fast-moving environments where the path forward isn’t laid out for youUnderstanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores)Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation)3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systemsFamiliarity with at least one deep learning framework (PyTorch, JAX, TensorFlow)Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision servingML compilers and framework internals: PyTorch internals, torch.compile, custom operatorsProfiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysisContainer orchestration: Kubernetes, GPU scheduling, autoscaling inference workloadsDistributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism

#J-18808-Ljbffr

AI Inference Engineer (Member of Technical Staff) employer: Perplexity AI

Perplexity is an exceptional employer, offering a dynamic work environment in the heart of London where innovation thrives. With a strong focus on employee growth and collaboration, team members are encouraged to push the boundaries of search technology while enjoying competitive compensation and equity options. The company's rapid expansion and backing from leading investors create unique opportunities for meaningful contributions and career advancement.

Perplexity AI

Contact Details:

Perplexity AI Recruitment Team