Inference Systems Engineer (GPU/Cluster Scaling)

Inference Systems Engineer (GPU/Cluster Scaling)

Full-Time On-site
AItoolnavio

Luma is seeking a systems engineer to own large-scale model serving and inference pipelines. You will integrate new architectures into the inference engine, scale deployments across thousands of machines, and keep GPU fleets busy while meeting internal SLOs.

You’ll work on scheduling, fleet management, and reliable deployment pipelines across clusters, with a strong emphasis on efficiency and uptime. This is the systems side of ML, not pure modeling.

#J-18808-Ljbffr

AItoolnavio

Contact Details:

AItoolnavio Recruitment Team