Principal Technical Program Manager

MicrosoftWashington, DC
$142,800 - $304,200Hybrid

About The Position

Cloud Operations + Innovation (CO+I) builds, operates, and evolves the infrastructure behind the Microsoft Cloud. The CO+I Artificial Intelligence (AI) Delivery organization is responsible for enabling Microsoft’s next generation of hyperscale AI supercomputing capacity. The AI Network Acceleration team focuses on making AI capacity delivery predictable, repeatable, and fast at global scale. This Principal Technical Program Manager role is central to that mission, driving design validation and readiness assessment for AI data center network infrastructure, with a deep focus on GPU platforms, backend networks, and high-performance fabrics. The role involves co-owning backend network readiness with platform and engineering partners to ensure AI network designs are technically sound, deployable, scalable, and operationally viable. This is a high-impact role at the intersection of architecture, execution, and acceleration, directly influencing the speed at which AI capacity reaches production and shaping the global deployment and scaling of GPU clusters and AI networks.

Requirements

  • Bachelor's Degree AND 6+ years experience in Technical Program Management, AI Data center Infrastructure engineering, Cloud or data center platform delivery OR equivalent experience.
  • 3+ years of experience managing cross-functional and/or cross-team projects.
  • Ability to meet Microsoft, customer and/or government security screening requirements.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Nice To Haves

  • Strong technical judgment with the ability to engage credibly with engineering partners.
  • 6+ years’ experience in AI data center infrastructure, GPU platforms, or hyperscale networking.
  • Experience with one or more of the following: Backend fabrics using InfiniBand and/or advanced technologies, including optics; Large-scale GPU clusters (training and/or inference); High-density rack architectures, optics, and cabling systems.
  • Strong understanding of: AI workload communication patterns; Backend network scale limits and failure modes.
  • Experience reducing deployment cycle time through: Design standardization; Automation and tooling; Parallel bring-up and validation models.
  • Exposure to AI platform bring-up, capacity activation, or production readiness reviews.

Responsibilities

  • Turn record-breaking AI infrastructure acceleration into the global operating standard.
  • Lead Global AI Infrastructure Deployment Acceleration: Own the development, training workforce and scaling of a global AI network and infrastructure acceleration operating model spanning Commercial Cloud, new AI regions, new data centers, colo expansions, and large-scale GPU deployments.
  • Translate successful deployment practices into repeatable architectures, operating mechanisms, technical decision frameworks, runbooks and deployment patterns that can be reused globally.
  • Identify where AI infrastructure delivery remains dependent on manual coordination, repeated escalation, organizational handoffs or individual expertise and systematically eliminate those dependencies.
  • Provide Deep Technical Program Leadership: Develop strong end-to-end understanding of the AI infrastructure deployment lifecycle and its critical dependencies across areas including: GPU infrastructure and large-scale AI clusters, Data-center network architecture, High-performance Ethernet and InfiniBand environments, Front-end and back-end network infrastructure, Regional and WAN connectivity, RNG and IDF readiness, Data-center fit-out and infrastructure readiness, Capacity planning and deployment sequencing, Facility/network dependencies, Infrastructure validation and acceptance, High-speed optical technologies and evolving network architectures, Cloud and distributed infrastructure systems.
  • Challenge assumptions, recognize architectural and execution risks, facilitate technical decisions, and translate engineering complexity into actionable execution plans.

Benefits

  • Relocation support will be provided
  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service