Senior Observability & Telemetry Engineer for GPU Cloud

Senior Observability & Telemetry Engineer for GPU Cloud

Full-Time 60000 - 80000 Β£ / year (est.) No working from home possible
Submer

At a Glance

  • Tasks: Design and implement an observability platform for large-scale GPU cloud infrastructure.
  • Company: Join Submer, a leader in innovative tech solutions.
  • Benefits: Hybrid work environment and great career growth opportunities.
  • Other info: Collaborate with diverse engineering teams in a dynamic setting.
  • Why this job: Make a real impact on cutting-edge AI workloads and distributed systems.
  • Qualifications: Experience in monitoring AI workloads and strong programming skills in Go or Python.

The predicted salary is between 60000 - 80000 Β£ per year.

Submer is looking for a senior engineer to design and implement an observability platform tailored for large-scale GPU cloud infrastructure and edge deployments. The role involves building telemetry pipelines, monitoring system behaviours, and collaborating with various engineering teams.

The ideal candidate should have experience in monitoring AI workloads, a strong programming background in Go or Python, and a deep understanding of distributed systems observability.

Submer offers a hybrid-friendly work environment and opportunities for career growth.

Senior Observability & Telemetry Engineer for GPU Cloud employer: Submer

Rubix is an exceptional employer, offering a unique opportunity for the Vice President, Legal to shape and lead a critical function within a rapidly growing global AI and digital infrastructure platform. With a commitment to innovation and sustainability, employees benefit from a dynamic work culture that fosters collaboration and inclusivity, alongside competitive compensation and significant opportunities for professional growth in a high-impact environment.

Submer

Contact Details:

Submer Recruitment Team

We think you need these skills to ace Senior Observability & Telemetry Engineer for GPU Cloud

Observability Platform Design
Telemetry Pipeline Development
Monitoring System Behaviours
AI Workload Monitoring
Programming in Go
Programming in Python
Distributed Systems Observability