Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data‑centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations
Initial 3 month contract with the possibility to extend.
Role Summary:
We are looking for Platform Engineer (HPC & AI) who canassistin shaping our new Platform team, this role will be customer facing, involve technical troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.
Responsibilities:
- Designing, deploying, and managing large‑scale HPC and GPU‑accelerated clusters, including NVIDIAbasedcomputeenvironments.
- Implementing and administering HPC scheduling and resource‑management systems (e.g.,Slurm), including GPU partitioning, workload scheduling, and capacity planning.
- Architecting and optimising InfiniBand and Ethernet network topologies.
- Ensuring high availability and resilience through failover strategies, planned maintenance coordination, and proactive risk mitigation.
- Automating provisioning, configuration, monitoring, and operational workflows across multi‑vendor HPC hardware and software stacks.
- Monitoring real‑time performance and leading troubleshooting efforts across compute, storage, interconnect, drivers, and node failures, engaging vendor support for critical issues.
- Incident response: node failure management, network issues, driver issues, troubleshooting common issues and then working with vendor support to resolve any critical issues.
- Security and access control: Manage user permissions, RBAC, security hardening, data protection.
Required Skills & Experience:
- Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms.
- System administration experience with OS's like RHEL/CentOS, Ubuntu, tuning Linux kernel.
- Proficiencywith Ansible, Nvidia and CUDA toolkits,Kubernetesand container orchestration.
- Understanding of automation,monitoringand security with GPU as a service.
- Extensive experience in system engineering, platformoperationsor SRE.
- Experience with GPU resource allocation (across instances, GPUs count and time).
- Advanced networking skills with High performance networking,troubleshootingand fine tuning.
- Familiarity with cloud-based platforms, APIs, and distributed systems.
- Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics).
- Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk).
- Excellent communication skills to interface with both customers and internal / vendor teams.
- Good understanding of tools requirements for ML engineers and data scientists, and how to optimise the experience.
Why Join Era4:
You’llbe joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation companyoperatesat scale.
Diversity & Inclusion :
Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Technology
United Kingdom: (Occasional office visit required)
#J-18808-Ljbffr
Platform Engineer - Contract employer: Carbon3ai Limited.
Era4 is an exceptional employer, offering a dynamic work environment in Bristol where innovation meets sustainability. As a Service Desk Analyst, you'll play a crucial role in supporting cutting-edge AI infrastructure while enjoying opportunities for professional growth and development within a mission-driven start-up culture that values diversity and inclusion.