Callosum, based in London, is seeking an experienced infrastructure-focused engineer to own end-to-end performance for its heterogeneous inference platforms. The role spans KV cache strategies, batching, memory management and multi-node scheduling to drive speed and efficiency across a growing model portfolio.
You will design, profile, and optimise GPU kernels and networking paths, build tooling for visibility, and collaborate across the stack to scale hardware and models.
#J-18808-Ljbffr