Inferact is seeking an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, and benchmarking infrastructure using ROCm, HIP, Triton, CK, and AITER.
You will work at the boundary of inference systems, kernels, compilers, and hardware, improving attention, GEMM, sampling, KV cache, and other communication-heavy paths to deliver fast, scalable inference and maintainable backend.
#J-18808-Ljbffr
Contact Details:
INFERACT SINGAPORE PTE. LTD. Recruitment Team