Callosum is building the next generation of AI infrastructure in London. We own the production system through end-to-end health, SLOs, observability, and incident response, across heterogeneous compute backends.
You will set the direction for reliability and partner with platform and hardware teams to keep systems resilient at scale. You will drive capacity planning, on-call readiness, and deployment discipline in a fast-growing environment, ensuring production-grade operations for real traffic
#J-18808-Ljbffr