At a Glance
- Tasks: Architect and manage cloud infrastructure for MLOps and HPC workloads globally.
- Company: Join a forward-thinking tech company focused on innovation and collaboration.
- Benefits: Relocation support, flexible working options, and a commitment to diversity.
- Other info: Opportunity for career growth in a dynamic and supportive environment.
- Why this job: Lead technical projects that enhance system reliability and make a global impact.
- Qualifications: Expertise in Infrastructure as Code and cloud-native architectures required.
The predicted salary is between 60000 - 80000 £ per year.
Responsibilities
- Architect Infrastructure as Code (Terraform, Pulumi, Cloud Formation) to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.
- Design for resilience, building disaster recovery and failover plans with auto‑scaling and load balancing to keep critical systems available worldwide.
- Strengthen reliability through chaos engineering experiments that validate systems and surface weaknesses before incidents.
- Build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.
- Provide technical leadership to a team of engineers, fostering collaboration, innovation, and continuous improvement.
- Partner across teams to align infrastructure with ML and HPC needs and advance operational maturity through SLAs, SLOs, SLIs, and error budgets.
Qualifications
- Deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or Cloud Formation in AWS, Azure, or GCP for MLOps and HPC workloads.
- Understanding of cloud‑native and on‑prem architectures including autoscaling, serverless, and multi‑region deployments, and hands‑on experience with Docker, Kubernetes, and Kubeflow.
- Expertise in automation scripting confidently in Python, Bash, or Go, with a strong grasp of GPU‑accelerated computing and HPC workload scaling.
- Leadership skills through influence, with strong communication, mentoring, and problem‑solving abilities.
- Degree in Computer Science or related technical field, or equivalent experience in software and site reliability engineering.
- Preferred experience with distributed ML frameworks such as Horovod or Tensor Flow Distributed, familiarity with data engineering pipelines such as Apache Airflow or Spark, and knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.
Benefits
- Relocation benefits available for this position.
All qualified applicants will receive consideration for employment without regard to race, religion or belief, sex, gender reassignment, sexual orientation, marriage and civil partnership, pregnancy and maternity, disability or age.
We recognise the importance of flexible working and will review all applicants’ requests with care.
#J-18808-Ljbffr
Principal/Senior Site Reliability Engineer in Welwyn employer: F. Hoffmann-La Roche AG
Roche is an exceptional employer, offering a dynamic work environment in Welwyn that fosters innovation and collaboration within the Supply Chain and Engineering sectors. Employees benefit from a strong commitment to continuous development, inclusive culture, and opportunities for professional growth, all while being part of a high-performing team dedicated to delivering impactful procurement solutions. With a focus on psychological safety and community building, Roche empowers its workforce to experiment, learn, and celebrate successes together.