Senior Technical Program Manager, Cluster Operations & Quota Management

MicrosoftMountain View, CA
$142,800 - $304,200

About The Position

At Microsoft AI, compute is the foundation everything else is built on: every frontier training run, every eval, and every inference workload depends on our GPU fleet. GPUs are our scarcest and most valuable resource. Leadership sets how that fleet is allocated; this role owns the entire system that makes those allocation decisions real - moving, provisioning, and validating quota across a constrained pool as fast and as cleanly as possible. We are looking for a Technical Program Manager with deep, hands-on experience in cluster operations and compute capacity management to own quota and cluster operations execution end to end at MAI. You will own the entire system that turns those allocation decisions into usable capacity - reliably, at speed, and at growing scale. This is a high-agency, service-oriented role for someone who is relentless on detail and follow-through, never lets anything drop, and likes turning a fragmented, high-stakes process into a clean, scalable machine.

Requirements

  • Significant experience in technical program management within infrastructure, platform engineering, or other compute-intensive environments.
  • A record of proactively improving operating models: spotting process failures, convening the right people across organizations, and driving automation and streamlining without waiting to be asked.
  • High agency and a service-oriented mindset, comfortable relentlessly chasing and pushing across teams, without formal authority, to get things done fast.

Responsibilities

  • Own the end-to-end operating system for quota and cluster execution: the process, playbooks, tooling agenda, and cross-company coordination that turn allocation decisions into usable capacity
  • Execute approved quota and capacity allocations across MAI's squads end to end: prepare and sequence changes in advance, run cluster moves and cycle cutovers cleanly, and drive each one to completion across every partner team.
  • Own the full chain to usable capacity, not just configured quota: coordinate provisioning and environment readiness with infrastructure, platform, HPC, and vendor teams, and validate identity, access, and utilization so researchers are productive from day one.
  • Relentlessly chase dependencies, surface blockers early, and unblock them, keeping a crisp, auditable status on every change so nothing is dropped or misreported.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service