At Microsoft AI, compute is the foundation everything else is built on: every frontier training run, every eval, and every inference workload depends on our GPU fleet. GPUs are our scarcest and most valuable resource. Leadership sets how that fleet is allocated; this role owns the entire system that makes those allocation decisions real - moving, provisioning, and validating quota across a constrained pool as fast and as cleanly as possible. We are looking for a Technical Program Manager with deep, hands-on experience in cluster operations and compute capacity management to own quota and cluster operations execution end to end at MAI. You will own the entire system that turns those allocation decisions into usable capacity - reliably, at speed, and at growing scale. This is a high-agency, service-oriented role for someone who is relentless on detail and follow-through, never lets anything drop, and likes turning a fragmented, high-stakes process into a clean, scalable machine.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed