Together AI is seeking a Staff Software Engineer to design a Kubernetes-native control plane that provisions and runs a GPU inference fleet across London and Amsterdam. You’ll build a manifest-driven API, contribute self-service tooling, and optimize scheduling, defragmentation, and capacity balancing to push utilization while preserving latency.
You’ll own the provisioning lifecycle, build reliable pipelines, and partner with the ML platform team to encode cluster shapes as first-class
#J-18808-Ljbffr