Callosum is seeking a seasoned ML inference engineer to push heterogeneous hardware orchestration beyond GPUs. You will work on SGLang and vLLM, extending them to run efficiently across diverse accelerators, focusing on scheduling, memory and execution.
Based in London, you will collaborate with an Accelerator Systems Software engineer, contribute upstream, and maintain internal forks while scaling production inference on mixed hardware.
#J-18808-Ljbffr
Hardware-Aware Inference Engine Engineer in London employer: UNCOVER
Join a prestigious international law firm in London, where you will be part of a high-performing legal team dedicated to excellence and innovation. The firm offers a collaborative work culture that fosters professional growth, providing ample opportunities for career advancement and skill development. With a focus on supplier agreements and commercial technology contracts, this role not only promises meaningful work but also the chance to make a significant impact within a globally recognised organisation.