Jobgether is seeking a Senior Site Reliability Engineer for a tokens-based inference platform in the UK. You will own reliability, performance, and observability for large-scale GPU-heavy workloads, designing telemetry, monitoring, and automation to ensure high availability.
You'll collaborate with software and infra teams to build self-healing systems, improve SLOs, and drive incident response with robust post-mortems.
#J-18808-Ljbffr
Senior SRE - AI Inference Platform (GPU/Kubernetes) employer: Jobgether
As a Senior ML Engineer at our innovative healthcare-focused company in the UK, you will be part of a dynamic team dedicated to transforming clinical environments through cutting-edge technology. We offer a collaborative work culture that fosters creativity and professional growth, alongside competitive benefits and opportunities for continuous learning in the rapidly evolving field of machine learning. Join us to make a meaningful impact on healthcare while enjoying a supportive environment that values your contributions.