Technical Program Manager - Cluster Build

Advanced Micro Devices, IncUNAVAILABLE, Texas

About The Position

In this high-profile role, you will serve as the technical program leader for large-scale compute cluster bring-up, driving execution from planning through delivery. You will lead cross-functional hardware, infrastructure, software, firmware, and validation teams to deliver operational clusters on schedule while coordinating issue resolution across the organization. You will drive the processes, dependencies, and execution rhythms that transform build plans into fully operational environments. Success in this role requires strong analytical, problem-solving, and risk-management skills, as well as the ability to build effective relationships and drive alignment across a complex, fast-moving engineering organization. You must be self-directed and comfortable operating in dynamic environments with significant cross-functional dependencies. As the Cluster Build Program Manager, you are the accountable leader for cluster bring-up from concept through delivery. You partner with engineering teams to define build strategies, assess technical feasibility, sequence dependencies, and drive execution against program objectives. You coordinate the teams responsible for the build, maintain alignment across stakeholders, and remove obstacles that threaten delivery. When issues arise during bring-up, you lead cross-functional triage, coordinate root-cause investigations, drive corrective actions, and ensure resolution through closure. The ideal candidate has a proven track record of delivering complex hardware and software programs, thrives in matrixed organizations, and consistently achieves predictable, on-time delivery of business-critical infrastructure.

Requirements

  • Bachelor's or Master's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field (or equivalent practical experience).
  • Formal project management education and PMP, Scrum Master, or equivalent certification/training.

Nice To Haves

  • Experience with AI/HPC Cluster Builds.
  • Detail-oriented and self-motivated, with a strong sense of accountability and pride in delivering results through cross-functional teams.
  • Proven ability to lead and coordinate cross-functional technical teams through complex, multi-stage cluster bring-up, integration, or infrastructure deployment efforts.
  • Comfortable driving cross-functional debug, triage, and issue resolution activities while guiding teams to root cause and corrective action.
  • Ability to quickly grasp technical concepts and convert meeting discussions into clear, actionable outcomes.
  • Excellent organizational, analytical, problem-solving, presentation, written, and verbal communication skills.
  • Results-oriented, highly organized, and capable of managing multiple priorities in fast-paced environments.
  • Proven ability to build productive relationships and collaborate effectively across organizations while operating as a self-starter.
  • Demonstrated horizontal leadership and matrix-management experience, with the ability to drive execution without direct authority.
  • Experience working effectively in high-pressure environments with competing priorities and aggressive timelines.
  • Strong executive communication skills, including program reviews, status reporting, and escalation management.
  • Demonstrated ability to collaborate across Engineering, Product Management, Business Units, and Program Management organizations to solve complex problems and mitigate risk.
  • People-management experience is desirable.
  • Experience working within Agile development environments.
  • Strong proficiency with project and productivity tools, including Microsoft Project, Jira, Confluence, and Microsoft Office Suite applications.

Responsibilities

  • Lead cluster bring-up end to end—from concept and planning through delivery and operational handoff.
  • Direct and coordinate cross-functional teams across hardware, infrastructure, software, firmware, validation, and site/datacenter operations.
  • Assess and scope the technical feasibility of cluster bring-up efforts by partnering with software and hardware architects, technical leads, and engineering teams to clarify requirements and dependencies.
  • Develop integrated program schedules that account for hardware, firmware, software, infrastructure, and bring-up dependencies. Align program milestones and delivery objectives.
  • Lead debug and issue resolution during bring-up by triaging problems, convening the appropriate subject-matter experts, coordinating root-cause investigations, and tracking corrective actions through closure.
  • Build, maintain, and execute comprehensive program plans, including plans of record, statements of work, schedules, budgets, and performance and quality KPIs.
  • Provide clear executive and operational visibility into bring-up status while facilitating dependency management and communication across teams.
  • Anticipate risks throughout the development of the lifecycle, identify mitigation strategies, and proactively address issues to protect schedules and optimize engineering resources.
  • Partner with core and execution teams to identify areas requiring escalation or corrective action and drive resolution.
  • Collect, analyze, organize, and publish bring-up performance data through dashboards, metrics, and recurring status reviews.
  • Drive continuous improvement of bring-up processes, tools, and methodologies to increase throughput, repeatability, and operational efficiency.
  • Work effectively with senior leaders across Engineering, Product Management, Customer Engineering, and business organizations to ensure alignment on priorities, risks, and program objectives.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service