Inferact Singapore PTE. LTD. is seeking a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs.
You will build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware. You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU
#J-18808-Ljbffr
Contact Details:
INFERACT SINGAPORE PTE. LTD. Recruitment Team