- CoreWeave is looking for an Engineering Manager to lead a team building and operating Kubernetes infrastructure on bare metal
- This team sits close to the core of our platform and is responsible for the reliability, scalability, and operational excellence of the systems that power high-performance AI and ML workloads
- You will lead engineers working on cluster lifecycle, platform reliability, infrastructure automation, and the operational systems that make Kubernetes run predictably at scale on dedicated hardware
- This is a hands-on leadership role for someone who can grow engineers, improve execution, and partner deeply with platform, networking, compute, and product teams
- The right person understands what it takes to run Kubernetes in demanding production environments and can turn that understanding into a high-functioning team, clear priorities, and durable engineering systems
- As the Engineering Manager for Kubernetes Infrastructure, you will lead a team responsible for the core infrastructure and operational foundations behind Kubernetes running directly on bare metal
- At CoreWeave, this platform is built for high-performance computing workloads and gives customers direct access to dedicated hardware, without virtualization overhead
- That makes reliability, performance, observability, and safe lifecycle management especially important
- You will be responsible for team execution, technical direction in partnership with senior engineers, and building strong operating mechanisms around delivery, quality, and incident response
- You will help the team scale its impact while mentoring engineers, supporting hiring, and strengthening cross-functional collaboration
- Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal
- Set clear goals, priorities, and execution plans for the team, and ensure reliable delivery against them
- Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation
- Improve the team’s operational excellence across incident response, on-call health, root-cause analysis, and service ownership
- Drive engineering best practices for safe change management, testing, rollout quality, and production readiness
- Support the design and operation of platform capabilities for provisioning, patching, upgrades, scaling, and troubleshooting of Kubernetes clusters
- Build strong cross-functional relationships with compute, networking, storage, security, and product stakeholders
- Hire, coach, and develop engineers while creating a high-accountability, high-trust team culture
- Establish and improve mechanisms for planning, prioritization, execution tracking, and continuous improvement
- Help translate complex platform and infrastructure work into clear business and customer value
- In the first 90 days
- Build trust with the team and key partner organizations
- Assess team health, role clarity, roadmap risks, and operational pain points
- Establish or tighten core operating rhythms for planning, execution, and incident follow-up
- Create a clear view of the highest-value reliability and scalability opportunities
- In the first 6 months
- Improve predictability of team execution and service ownership
- Raise the quality bar for change management, rollout safety, and operational readiness
- Strengthen hiring and development plans for the team
- Drive measurable improvements in one or more areas such as cluster reliability, upgrade safety, provisioning speed, or observability
- In the first 12 months
- Build a strong, durable team with clear ownership and healthy operating mechanisms
- Deliver meaningful infrastructure improvements that increase platform reliability, scalability, and maintainability
- Be recognized as a trusted cross-functional leader for Kubernetes infrastructure on bare metal
Strong written and verbal communication, including the ability to explain technical trade-offs and priorities clearlyAbility to coach engineers at different levels and create clarity in ambiguous or fast-scaling environmentsStrong technical depth in Kubernetes, distributed systems, and production infrastructureStrong partnership skills across engineering, product, and operations functionsFamiliarity with cluster lifecycle management, including provisioning, upgrades, node operations, observability, and reliability engineeringExperience managing an infrastructure, platform, or SRE-oriented engineering teamExperience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systemsTrack record of improving team execution, engineering quality, and operational maturityExperience leading incident response cultures and driving follow-through on reliability improvementsThis position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C. 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing processExperience with GPU-heavy, HPC, or ML infrastructure environmentsFamiliarity with Kubernetes networking, storage, and security primitives in productionExperience with bare-metal infrastructure, server lifecycle operations, or low-level systems troubleshootingExperience building internal platform products used by other engineering teams or external customersExperience with infrastructure automation using tools such as Go, Python, controllers/operators, or configuration management systems
#J-18808-Ljbffr
Engineering Manager (Kubernetes Infrastructure, Bare Metal) in Livingston employer: CoreWeave
CoreWeave is an exceptional employer that thrives on innovation and collaboration, making it an ideal place for the Senior Community Affairs & Partnerships Manager to make a meaningful impact. With a strong emphasis on employee growth, a dynamic work culture, and comprehensive benefits including flexible PTO and tuition reimbursement, CoreWeave fosters an environment where creativity and independent thinking are encouraged. Located in a fast-paced industry, employees are surrounded by top talent and have the opportunity to contribute to groundbreaking advancements in AI technology while building trusted relationships within the community.