Capacity Operations Manager

BasetenSan Francisco, CA
$225,000 - $235,000

About The Position

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.

Requirements

  • 5 to 10+ years within infrastructure working within the compute lifecycle to maximize functional compute, ideally in a hyperscale, cloud, or large-scale compute environment.
  • Direct experience managing GPU, server, or data center hardware supplier relationships. You understand fleet health, RMA processes, and how contracted capacity differs from delivered capacity.
  • Highly analytical. You should be comfortable pulling your own data, building your own reports, and generating insights without waiting on someone else to hand you a dashboard.
  • Comfortable with ambiguity. Part of the job is figuring out what should exist and building it.
  • Strong cross-functional collaboration skills. You'll work closely with finance, infrastructure/engineering, legal, and security on a regular basis.

Nice To Haves

  • Experience at a hyperscaler, neo cloud provider, or AI infrastructure company
  • Familiarity with GPU hardware lifecycles (NVIDIA H100/H200/GB200 class systems), power/thermal constraints, and supply chain dynamics for compute.
  • Experience running formal supplier corrective actions.

Responsibilities

  • Drive suppliers to keep the maximum amount of the GPU fleet online and healthy.
  • Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs.
  • Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier.
  • SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short.
  • Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability.

Benefits

  • Competitive compensation, including meaningful equity
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service