Senior Manager, Technical Support Engineering - Cloud

CoreWeaveSunnyvale, CA
$198,000 - $264,000Onsite

About The Position

The Technical Support Engineering - Infrastructure team sits within CoreWeave's Global Field Organization (GFO) and is the front line for every customer running AI and HPC workloads on our platform. Operating 24/7/365, the team supports the Kubernetes-powered infrastructure behind the AI revolution: GPU compute, high-performance networking and storage, Slurm/HPC clusters, and the large-scale, mission-critical training workloads that run on them. We operate on a Direct-to-Expert model. Instead of routing customers through generic support tiers, we get the right expert on the problem fast while a single owner stays accountable for the issue end-to-end. That means fewer handoffs, clearer ownership, and a high-touch experience that our customers feel on every ticket. The team is the connective tissue between our customers and CoreWeave's engineering organization. We triage and resolve deep technical issues directly, coordinate with Product Engineering, Specialist Field Engineers, and domain specialists when needed, and feed what we learn in the field straight back into the product roadmap. As the Senior Manager of Technical Support Engineering - Infrastructure, you'll own, scale, and continuously improve CoreWeave's infrastructure support function — a large, globally distributed 24/7/365 organization of skilled engineers who resolve our customers' most complex technical challenges with deep expertise, efficiency, and empathy. High-touch, expert-led support is one of CoreWeave's clearest differentiators, which gives this role prominent visibility and influence across the company. You'll operate at the department level, setting the technical and operational bar for the entire function and designing the operating systems (coverage/staffing model, Direct-to-Expert partnerships, quality program, metrics) that let support scale reliably. You'll lead with empathy and invest in each person's growth, while building the structure and processes that let the team scale with CoreWeave's hyper-growth without compromising the culture we care deeply about.

Requirements

  • 5+ years of people-leadership experience and 8+ years total in technical support/operations, including running a 24/7 support function at scale in a cloud operations environment.
  • A strong background in Linux, containerization technologies, and Kubernetes, and you understand virtualization and cloud computing concepts.
  • Leadership & Communication: proven ability to lead through senior talent and set direction for a function, with executive-level communication skills.
  • Strategic & Operational Planning: you build operating systems, plans, and metrics that scale a function beyond what any one person can hold.
  • Problem-Solving & Adaptability: robust problem-solving skills and adaptability at organizational scale, in a fast-paced, hyper-growth environment.
  • Program Management: experience with program-management tools and methodologies.
  • You lead with empathy and aren't afraid to get your hands dirty. You do the work alongside your direct reports and model the standard you set.
  • You're energized by leadership excellence and talent development: diligent performance management, coaching, and growing each engineer according to their individual needs.
  • You've built enablement and quality programs that scaled across a function or multiple teams — and can show the measurable improvement they drove.
  • You've designed the operating model for a support org, including coverage/staffing model, escalation boundaries, SLOs, and evolved it as the org scaled.
  • You're a calm, clear communicator with customers and executives during critical incidents, and you resolve conflicts effectively across teams and organizational boundaries.
  • You think in systems and multi-quarter plans. You've defined KPI/SLO frameworks, reported to senior leadership, and owned capacity and growth planning for a function, not just a single team.
  • You've led a globally distributed team across time zones.

Nice To Haves

  • Experience at a hyperscaler or cloud infrastructure provider is a strong plus.
  • Experience supporting AI/ML, HPC, or GPU-accelerated workloads at scale.
  • Hands-on Kubernetes operations experience (CKA certification a plus).
  • Familiarity with Slurm/SUNK, RDMA networking, distributed storage, and observability tooling such as Grafana.
  • You have some experience with infrastructure as it relates to Data Center Operations.
  • You're an expert in what it takes to be an excellent leader and foster an environment in which people are excited and inspired to participate.
  • You love to dive into problems, test for solutions, and enjoy engaging with customers.
  • You're excited and curious about AI.

Responsibilities

  • Own the strategy, health, and performance of the entire 24/7/365 infrastructure support function, helping scale coverage and capability across regions and domains as CoreWeave grows.
  • Own talent acquisition and retention by hiring, onboarding, and developing engineers through diligent performance management and coaching tailored to each individual's needs.
  • Stay hands-on: dig into complex, customer-impacting issues alongside your team and serve as a senior technical escalation point, ensuring the highest quality of support.
  • Build and facilitate the enablement and career-development frameworks for onboarding, technical training, and progression paths that raise capability across the whole function and create a repeatable path for career growth.
  • Implement quality assurance measures, including ticket reviews and best-practice playbooks, that raise the bar on resolution speed, accuracy, and consistency.
  • Own the support operating model for the function including roles and responsibilities, escalation boundaries, coverage/staffing model, and the response and resolution targets (SLOs) that keep every customer's experience consistently excellent as volume scales.
  • Lead customer communication during critical incidents and resolve conflicts with clarity, composure, and empathy.
  • Track and report on KPIs focused on team performance and customer satisfaction, and own the strategic planning for the team's growth and scalability.
  • Own the cross-functional interface between support and Product Engineering, Specialist Field Engineers, and domain teams, building the alignment mechanisms and feedback loops that make the Direct-to-Expert model work at scale for escalations, live incidents, and customer needs.
  • Champion the voice of the customer, turning recurring support patterns into product, tooling, and process improvements.
  • Help set the multi-quarter vision and operating plan for infrastructure support, and represent the function in company-level planning, headcount, and prioritization discussions.

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service