Sr. Director of Engineering, Data Center Infrastructure

Advanced Micro Devices, Inc•Seattle, WA

About The Position

AMD is seeking a visionary Senior Director of AI Data Center Infrastructure Engineering to lead the strategy, architecture, deployment, and operation of next-generation AI data centers and large-scale compute environments. This highly visible leadership role will be responsible for defining AMD's AI infrastructure architecture across rack-scale and cluster-scale deployments, spanning facility infrastructure, compute systems, networking, software operations, and reliability engineering. The successful candidate will lead a world-class organization of engineers and architects responsible for building scalable AI infrastructure that supports AMD's internal development, validation, and customer enablement initiatives. This position requires a unique combination of executive leadership, systems thinking, and deep technical expertise across data center architecture, networking, distributed systems, cluster operations, and large-scale infrastructure deployment.

Requirements

  • Experience designing, deploying, and operating large-scale data center environments.
  • Experience leading engineering organizations, including ownership of large multidisciplinary teams.
  • Demonstrated success building and scaling AI, cloud, hyperscale, HPC, or enterprise data center infrastructure.
  • Extensive experience with data center architecture, power delivery, thermal management, capacity planning, and facility operations.
  • Deep knowledge of AI infrastructure, server architectures, accelerator-based computing platforms, and rack-scale system design.
  • Proven expertise designing high-performance data center networks, including switching, routing, Ethernet fabrics, RDMA, InfiniBand, and next-generation interconnect technologies.
  • Experience leading compute cluster deployment, lifecycle management, and large-scale distributed system operations.
  • Strong understanding of Site Reliability Engineering (SRE), Infrastructure as Code (IaC), automation frameworks, monitoring, observability, and operational best practices.
  • Proven ability to lead technical organizations through significant technology transformations and rapid growth.
  • Experience influencing cross-functional roadmaps involving hardware, software, networking, and infrastructure teams.
  • Strong executive communication skills with demonstrated ability to influence senior leadership, customers, and external partners.
  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, Mechanical Engineering, or related technical field.

Nice To Haves

  • Master's degree preferred.
  • PhD considered a plus.

Responsibilities

  • Define AMD's end-to-end AI data center infrastructure strategy, including compute architecture, facility requirements, operational models, and scalability planning.
  • Lead architecture and deployment decisions for large-scale AI and HPC data center environments spanning thousands of servers and accelerators.
  • Drive infrastructure planning across power distribution, cooling strategies, space utilization, capacity forecasting, and operational readiness.
  • Analyze performance, efficiency, reliability, and total cost of ownership (TCO) tradeoffs to optimize infrastructure investments.
  • Partner with internal engineering organizations, facilities teams, and external vendors to deliver world-class AI infrastructure platforms.
  • Own architecture decisions spanning silicon, servers, rack-scale solutions, and full data center deployments.
  • Collaborate closely with hardware, software, networking, systems design, platform engineering, and product organizations to align infrastructure capabilities with AMD technology roadmaps.
  • Drive architecture reviews and technical decision-making across mechanical, electrical, thermal, firmware, software, and systems domains.
  • Establish scalable infrastructure standards and design methodologies for future AI platforms.
  • Lead the design, deployment, and lifecycle management of large-scale AI training and inference clusters.
  • Establish operational excellence through Site Reliability Engineering (SRE) principles, observability frameworks, incident management processes, and capacity planning.
  • Develop automation strategies that improve deployment speed, infrastructure efficiency, system reliability, and operational scalability.
  • Drive continuous improvement initiatives focused on availability, resilience, performance, and cost optimization.
  • Define reliability and service-level objectives (SLOs) for production infrastructure environments.
  • Define and optimize networking architectures supporting large-scale AI and HPC environments.
  • Lead strategy across network topology, high-speed fabrics, switching architectures, routing infrastructure, and accelerator interconnect technologies.
  • Evaluate and implement next-generation networking technologies to maximize cluster performance and scalability.
  • Champion software-defined networking (SDN), network automation, and observability solutions across global infrastructure deployments.
  • Partner with AMD product and ecosystem teams to influence future networking roadmaps.
  • Lead, mentor, and grow a global organization of 50-75 engineers, architects, and technical leaders.
  • Establish a high-performance engineering culture focused on innovation, accountability, collaboration, and execution excellence.
  • Develop organizational strategy, succession planning, workforce development, and talent acquisition initiatives.
  • Serve as a key technical advisor to executive leadership on AI infrastructure strategy, investment priorities, and industry trends.
  • Collaborate with strategic customers, system integrators, technology partners, and suppliers to drive alignment across the AI ecosystem.
  • Monitor competitive technologies, emerging market trends, and industry innovation to influence AMD's long-term infrastructure vision.

Benefits

  • AMD benefits at a glance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service