Responsibilities
- The AI/ML TPM team owns delivery and execution across CoreWeave’s AI/ML Platform Services organization
- The team partners closely with Product, Engineering, Research, Infrastructure, and Go-to-Market teams to deliver scalable, reliable, and high-performance platforms that support the full AI lifecycle
- AI/ML TPMs drive alignment and execution across highly technical, cross-functional teams to ensure the successful delivery of customer-facing infrastructure and platform capabilities used by researchers, engineers, and enterprise customers
- As a Technical Program Manager focused on inference, you will lead complex, cross-functional programs spanning inference platform delivery, customer onboarding, launch readiness, and runtime optimization
- The Inference team is responsible for building and operating highly scalable, reliable production inference services for both serverless and dedicated inference use cases
- The work spans platform reliability, operational excellence, customer onboarding, release validation, runtime performance, and the launch of new inference capabilities that help customers run models at scale with strong price-performance and low operational friction
- In this role, you will partner with engineering, product, infrastructure, and go-to-market teams to drive programs that improve how inference services are launched, onboarded, operated, and optimized
- This includes managing the execution of customer-facing onboarding programs, launch readiness for dedicated inference capabilities, and the platform improvements needed to support scale, reliability, and predictable delivery
- Drive end-to-end program management for inference platform initiatives spanning reliability, customer onboarding, launch readiness, and runtime optimization
- Lead cross-functional programs for customer onboarding across dedicated and serverless inference offerings, ensuring clear ownership, launch criteria, and readiness for strategic customer use cases
- Drive launch readiness for new inference capabilities by aligning teams around real customer outcomes, supportability, and end-to-end validation
- Partner with engineering and product to define and deliver roadmap outcomes for latency, throughput, uptime, operational quality, and price-performance
- Coordinate multi-team execution across platform, infrastructure, and customer-facing teams to deliver reliable and scalable inference services
- Build and operationalize success metrics, dashboards, launch gates, and review cadences to measure service reliability, onboarding readiness, efficiency, and quality across the inference stack
- Establish repeatable processes for release validation, performance regression tracking, launch management, and postmortem follow-through
- Help unify operational processes, support mechanisms, and execution visibility across inference deployment models and customer onboarding paths
- Create strong communication channels between Engineering, Product, Infrastructure, and Go-to-Market teams to align priorities and deliver predictable, high-impact outcomes
Requirements
Strong technical fluency in distributed inference systems, GPU compute, cloud-native architectures, and performance optimization
Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren’t a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk
- Understanding of customer onboarding for technical products, especially where platform capabilities, infrastructure readiness, and support processes must align for launch
- Bachelor’s degree in a technical field or equivalent practical experience
- Proven experience driving large-scale infrastructure or platform programs from concept to production in complex, cross-functional environments
- Excellent written and verbal communication skills, with the ability to align engineering, product, infrastructure, and customer-facing stakeholders around shared goals
- Experience with inference-serving systems, model onboarding workflows, rollout strategies, and observability tooling
- Familiarity with launch readiness, supportability, incident follow-through, and release validation for production infrastructure or platform services
- 8+ years of technical program management experience in distributed systems, cloud infrastructure, or AI/ML platform engineering
- Demonstrated success driving measurable improvements in reliability, performance, operational readiness, or customer delivery
- Experience operating in high-growth environments where roadmap execution, reliability expectations, and customer commitments must be managed in parallel
- You’re effective at creating clarity and momentum across ambiguous, fast-moving multi-team initiatives
- You love driving execution for complex, customer-facing AI infrastructure and platform programs
- You’re curious about how large-scale inference systems evolve across runtime performance, operational excellence, and customer onboarding
- You enjoy turning technically complex platform work into predictable execution and successful launches
#J-18808-Ljbffr
Technical Program Manager (Inference) employer: CoreWeave
CoreWeave is an exceptional employer that thrives on innovation and collaboration, making it an ideal place for the Senior Community Affairs & Partnerships Manager to make a meaningful impact. With a strong emphasis on employee growth, a dynamic work culture, and comprehensive benefits including flexible PTO and tuition reimbursement, CoreWeave fosters an environment where creativity and independent thinking are encouraged. Located in a fast-paced industry, employees are surrounded by top talent and have the opportunity to contribute to groundbreaking advancements in AI technology while building trusted relationships within the community.