Member of Technical Staff (AI Inference Engineer)

Member of Technical Staff (AI Inference Engineer)

Full-Time 76023 - 95029 Β£ / year (est.) No working from home possible
CVFine by Instrovate Technologies

At a Glance

  • Tasks: Join our team to build and optimise AI inference engines for cutting-edge models.
  • Company: Perplexity, a forward-thinking tech company in London.
  • Benefits: Competitive salary, equity options, and opportunities for professional growth.
  • Other info: Dynamic work environment with a focus on collaboration and continuous learning.
  • Why this job: Make an impact in AI while working with innovative technologies and talented peers.
  • Qualifications: 3+ years in software engineering with a focus on ML inference and GPU programming.

The predicted salary is between 76023 - 95029 Β£ per year.

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.

Responsibilities:

  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.

Who we're looking for:

  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.
  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
  • Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.

Nice-to-have:

  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.

Qualifications:

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Final offer amounts are determined by multiple factors including experience and expertise. Equity: In addition to the base salary, equity may be part of the total compensation package.

Member of Technical Staff (AI Inference Engineer) employer: CVFine by Instrovate Technologies

Perplexity is an exceptional employer for AI Inference Engineers, offering a dynamic work environment in London where innovation thrives. With a strong focus on employee growth, we provide opportunities to work with cutting-edge technologies and contribute to high-performance systems that impact real-world applications. Our collaborative culture encourages self-direction and creativity, ensuring that every team member can make meaningful contributions while enjoying competitive compensation and equity options.

CVFine by Instrovate Technologies

Contact Details:

CVFine by Instrovate Technologies Recruitment Team

We think you need these skills to ace Member of Technical Staff (AI Inference Engineer)

GPU Programming
CUDA
Rust
Python
CuTe DSL
Performance Optimisation
Distributed Systems