Engineering Manager, Kubernetes Infrastructure (Bare Metal)

CoreWeaveSunnyvale, CA
$182,000 - $242,000Onsite

About The Position

CoreWeave is seeking an Engineering Manager to lead a team focused on building and operating Kubernetes infrastructure on bare metal. This critical team is responsible for the reliability, scalability, and operational excellence of the systems powering high-performance AI and ML workloads. The role involves leading engineers in areas such as cluster lifecycle management, platform reliability, infrastructure automation, and the operational systems that ensure Kubernetes runs predictably at scale on dedicated hardware. This is a hands-on leadership position requiring the ability to grow engineers, enhance execution, and collaborate closely with platform, networking, compute, and product teams. The ideal candidate possesses a deep understanding of running Kubernetes in demanding production environments and can translate this knowledge into a high-functioning team, clear priorities, and robust engineering systems. As the Engineering Manager for Kubernetes Infrastructure, you will lead a team responsible for the core infrastructure and operational foundations of Kubernetes running directly on bare metal, optimized for high-performance computing workloads with direct customer access to dedicated hardware. Key responsibilities include team execution, technical direction, building strong operating mechanisms for delivery, quality, and incident response, scaling team impact, mentoring engineers, supporting hiring, and strengthening cross-functional collaboration. The role involves establishing clear goals, priorities, and execution plans, partnering on roadmaps for cluster lifecycle management, upgrades, reliability, observability, and automation, improving operational excellence in incident response and on-call health, driving engineering best practices for change management and production readiness, supporting the design and operation of platform capabilities, building cross-functional relationships, hiring and developing engineers, and establishing mechanisms for planning and continuous improvement. The manager will also translate complex infrastructure work into clear business and customer value.

Requirements

  • Experience managing an infrastructure, platform, or SRE-oriented engineering team
  • Strong technical depth in Kubernetes, distributed systems, and production infrastructure
  • Experience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systems
  • Familiarity with cluster lifecycle management, including provisioning, upgrades, node operations, observability, and reliability engineering
  • Track record of improving team execution, engineering quality, and operational maturity
  • Experience leading incident response cultures and driving follow-through on reliability improvements
  • Strong partnership skills across engineering, product, and operations functions
  • Ability to coach engineers at different levels and create clarity in ambiguous or fast-scaling environments
  • Strong written and verbal communication, including the ability to explain technical trade-offs and priorities clearly

Nice To Haves

  • Experience with GPU-heavy, HPC, or ML infrastructure environments
  • Experience with bare-metal infrastructure, server lifecycle operations, or low-level systems troubleshooting
  • Familiarity with Kubernetes networking, storage, and security primitives in production
  • Experience building internal platform products used by other engineering teams or external customers
  • Experience with infrastructure automation using tools such as Go, Python, controllers/operators, or configuration management systems

Responsibilities

  • Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal
  • Set clear goals, priorities, and execution plans for the team, and ensure reliable delivery against them
  • Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation
  • Improve the team's operational excellence across incident response, on-call health, root-cause analysis, and service ownership
  • Drive engineering best practices for safe change management, testing, rollout quality, and production readiness
  • Support the design and operation of platform capabilities for provisioning, patching, upgrades, scaling, and troubleshooting of Kubernetes clusters
  • Build strong cross-functional relationships with compute, networking, storage, security, and product stakeholders
  • Hire, coach, and develop engineers while creating a high-accountability, high-trust team culture
  • Establish and improve mechanisms for planning, prioritization, execution tracking, and continuous improvement
  • Help translate complex platform and infrastructure work into clear business and customer value

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service