AI Cluster Technical Program Manager – Validation, Debug & Agentic AI

Advanced Micro Devices, IncAustin, TX
Onsite

About The Position

Join AMD's Cluster Validation and Platform Readiness organization, a highly visible team at the forefront of next-generation AI infrastructure. This team sits at the intersection of data center GPUs, AI rack solutions, cluster deployments, validation, debug, and customer readiness activities. You will work alongside architects, engineers, and technical leaders to help AMD's latest AI platforms from initial rack bring-up through deployment readiness and customer adoption. The role offers exposure to cutting-edge AI technologies, large-scale data center systems, and the opportunity to influence how AMD delivers industry-leading AI solutions to customers around the world. You will help drive validation and debug programs, coordinate complex cross-functional efforts, develop executive-level communications, and champion modern AI productivity tools that improve efficiency across the program management organization. This is an ideal opportunity for someone who enjoys working in a fast-paced environment where technical depth, strategic thinking, and innovation come together.

Requirements

  • Self-driven technical program manager who thrives in a highly collaborative and rapidly evolving environment.
  • Comfortable navigating ambiguity, building relationships across diverse organizations, and influencing teams without direct authority.
  • Strong communication, organizational, and problem-solving skills.
  • Ability to effectively engage with engineers, architects, executives, and other stakeholders.
  • Ability to independently drive initiatives, proactively identify risks, communicate complex technical information clearly, and bring structure to challenging problems.
  • Passionate about continuous improvement, leveraging emerging AI technologies and productivity tools to streamline program execution and enhance team effectiveness.
  • Bachelor’s or master’s degree preferred in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Systems, Engineering Management, or a related technical field.
  • PMP, Scrum Master, Agile, or related project/program management certifications desired.

Nice To Haves

  • Technical Program Management experience supporting complex hardware, system, infrastructure, or platform development programs.
  • Experience with AI infrastructure, data center technologies, GPU-based platforms, rack-scale systems, or cluster environments.
  • Understanding of rack-level and cluster-level validation, testing, quality assurance, or platform readiness activities.
  • Experience managing program execution across EVT, DVT, PVT, or similar product development and validation phases.
  • Exposure to system-level debug, incident management, issue resolution, and root-cause investigation processes.
  • Ability to develop executive dashboards, metrics, status reporting, and data-driven decision-making frameworks.
  • Experience working with cross-functional engineering teams in fast-paced technical environments.
  • Strong knowledge of Jira, Confluence, Microsoft Office Suite, and related project management tools.
  • Familiarity with Microsoft Copilot, Agentic AI tools, workflow automation solutions, or AI-assisted productivity platforms.
  • Experience driving process improvements, operational excellence initiatives, or organization-wide efficiency programs.
  • Strong presentation, communication, stakeholder management, and matrix leadership skills.

Responsibilities

  • Lead validation, debug, and platform readiness programs for next-generation AI rack and cluster solutions.
  • Partner with engineering, architecture, product, and operations teams to define plans, schedules, milestones, and deliverables across multiple programs.
  • Drive validation activities through development lifecycle phases, tracking quality metrics, coverage, readiness, risks, and execution status.
  • Coordinate incident management, root-cause investigations, corrective actions, and recovery efforts during validation and deployment activities.
  • Develop executive-level dashboards, metrics, presentations, and status reports that communicate program health, risks, and progress.
  • Promote adoption of AI-enabled productivity tools to improve reporting, knowledge management, workflow automation, and overall program efficiency.
  • Within the first 90 days, build strong cross-functional relationships, gain an understanding of AMD's AI infrastructure ecosystem, and begin contributing to active validation and debug programs.
  • Within six months, independently lead program initiatives, drive execution plans, manage risks, and influence key technical and operational decisions.
  • Long term, become a trusted leader within the organization, helping shape validation strategies, improve operational processes, and accelerate readiness for future AI platform deployments.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service