At a Glance
- Tasks: Design and improve Kubernetes-based AI infrastructure for cutting-edge projects.
- Company: Join Era4, a mission-driven start-up transforming energy sites into modern data centres.
- Benefits: Enjoy competitive salary, flexible work options, and opportunities for professional growth.
- Other info: Be part of a diverse team committed to operational excellence and inclusivity.
- Why this job: Make a real impact in AI while working with innovative technologies and a passionate team.
- Qualifications: Hands-on experience with Kubernetes, GPU environments, and strong automation skills.
The predicted salary is between 36000 - 60000 £ per year.
Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data‑centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public‑sector organisations.
We are seeking a DevOps Engineer to design, operate, and continuously improve our Kubernetes‑based AI infrastructure. This role focuses on cloud‑native platform engineering, GPU‑accelerated workloads, reliability, automation, and customer enablement. You will play a key role in delivering a production‑grade AI platform that enables ML engineers, data scientists, and enterprise customers to build and run AI workloads at scale. You will be responsible for the reliability, scalability, and performance of our Kubernetes‑based GPU platforms. You will ensure our AI platform operates securely and efficiently while delivering an exceptional customer experience. This is a hands‑on platform engineering position focused on systems reliability, automation, and continuous improvement.
Key Responsibilities
- Kubernetes Platform Operations: Operate and evolve a production Kubernetes environment supporting GPU‑accelerated AI workloads. Manage cluster lifecycle (deployment, upgrades, scaling, resilience, multi‑node operations). Implement high availability, failover, and maintenance strategies to minimize disruption. Enable aaS capabilities and segmentation for multi‑tenant workloads. Infrastructure as code tooling and lifecycle. Network Overlays, Storage: Block, File and Object. Experience with Ansible, YAML, Terraform, Python, Jenkins and GitOps.
- GPU & AI Infrastructure Engineering: Manage NVIDIA GPU infrastructure within Kubernetes (device plugins, drivers, CUDA compatibility). Implement GPU partitioning and workload isolation strategies (e.g., MIG, quotas, namespaces). Monitor and optimize GPU utilization, workload efficiency, and cluster capacity. Support AI/ML training and inference workloads with performance tuning and best practices.
- Reliability, Monitoring & Automation: Design and maintain observability frameworks (metrics, logs, tracing). Implement proactive monitoring, alerting, and capacity planning. Lead incident response for platform‑level events and drive root cause analysis. Automate operational workflows and infrastructure provisioning (IaC, configuration management). Contribute to platform reliability engineering practices (SLOs, SLAs, error budgets).
- Security & Governance: Implement RBAC, network policies, and security hardening. Ensure secure multi‑tenant workload isolation. Maintain compliance, data protection, and access governance standards.
- Customer & Platform Enablement: Support customer lifecycle of onboarding, provisioning and operations. Provide guidance on workload configuration, scaling strategies, and best practices. Collaborate with engineering and vendor teams to resolve complex platform issues. Produce high‑quality technical documentation and operational playbooks.
Required Experience & Skills
- Strong hands‑on experience operating production Kubernetes clusters.
- Experience with GPU‑enabled Kubernetes environments.
- Solid Linux system administration, networking, storage and security skills.
- Experience with Infrastructure as Code and automation.
- Strong understanding of distributed systems, APIs, and cloud‑native architectures.
- Experience implementing monitoring and observability solutions (e.g., Prometheus, Grafana).
- Proven incident management and root cause analysis experience.
- Strong communication skills and ability to work cross‑functionally.
Desirable Experience
- Experience operating AI/HPC infrastructure.
- Deep understanding of Kubernetes scheduling, networking, and storage.
- Experience with high‑performance datacentre networking and tuning.
- Background in DevOps or Site Reliability Engineering (SRE).
Why Join Era4
You’ll be joining a mission‑driven start‑up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next‑generation company operates at scale.
Diversity & Inclusion
Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
DevOps Engineer employer: Carbon3.ai
Joining Era4 as a Control Room Operative in Ardley means becoming part of a mission-driven start-up that is at the forefront of building critical national infrastructure. With a strong emphasis on operational excellence, you will enjoy a dynamic work culture that promotes diversity and inclusion, alongside ample opportunities for professional growth and development within a supportive environment. This role not only offers real autonomy and visibility with leadership but also allows you to play a pivotal role in shaping the future of a next-generation company.
StudySmarter Expert Advice🤫
We think this is how you could land DevOps Engineer
✨Tip Number 1
Network, network, network! Get out there and connect with people in the industry. Attend meetups, webinars, or even local tech events. You never know who might have a lead on that perfect DevOps Engineer role!
✨Tip Number 2
Show off your skills! Create a GitHub profile showcasing your projects, especially those related to Kubernetes and AI infrastructure. This gives potential employers a taste of what you can do and sets you apart from the crowd.
✨Tip Number 3
Don’t just apply blindly! Tailor your approach for each company. Research Era4, understand their mission, and align your skills with their needs. When you apply through our website, make sure to highlight how you can contribute to their AI platform.
✨Tip Number 4
Prepare for interviews by brushing up on your technical knowledge and soft skills. Practice common DevOps scenarios and be ready to discuss your experience with Kubernetes, automation, and incident management. Confidence is key!
We think you need these skills to ace DevOps Engineer
Some tips for your application 🫡
Tailor Your CV:Make sure your CV is tailored to the DevOps Engineer role. Highlight your experience with Kubernetes, GPU infrastructure, and any relevant automation tools like Ansible or Terraform. We want to see how your skills match what we're looking for!
Craft a Compelling Cover Letter:Your cover letter is your chance to shine! Use it to explain why you're passionate about AI infrastructure and how you can contribute to our mission at Era4. Keep it concise but impactful – we love a good story!
Show Off Your Projects:If you've worked on any cool projects related to Kubernetes or AI workloads, don’t hesitate to mention them! We’re keen to see real-world applications of your skills, so share links or descriptions that showcase your work.
Apply Through Our Website:We encourage you to apply directly through our website. It’s the best way to ensure your application gets into the right hands. Plus, it shows us you’re genuinely interested in joining our team at Era4!
How to prepare for a job interview at Carbon3.ai
✨Know Your Kubernetes Inside Out
Make sure you brush up on your Kubernetes knowledge before the interview. Be ready to discuss your hands-on experience with production clusters, including deployment, upgrades, and scaling. They’ll want to hear about specific challenges you've faced and how you overcame them.
✨Showcase Your Automation Skills
Era4 is all about efficiency, so be prepared to talk about your experience with Infrastructure as Code tools like Terraform and Ansible. Bring examples of how you've automated workflows or improved system reliability in previous roles. This will show that you can contribute to their mission of continuous improvement.
✨Demonstrate Your Problem-Solving Abilities
Expect questions around incident management and root cause analysis. Prepare to share specific instances where you’ve led incident responses or resolved complex platform issues. Highlight your analytical skills and how you approach troubleshooting in a high-pressure environment.
✨Communicate Clearly and Collaboratively
Strong communication skills are key for this role, especially since you'll be working cross-functionally. Practice explaining technical concepts in simple terms, and think of examples where you’ve successfully collaborated with other teams. This will help demonstrate that you’re not just a tech whiz, but also a team player.