AI/HPC Data Center Rack Design Engr

Advanced Micro Devices, IncAustin, TX

About The Position

We are seeking a highly skilled systems engineer to design data center rack layouts for AI/HPC clusters. This role involves evaluating and selecting rack-level compute, storage, networking, power delivery, and cooling solutions to optimize performance and reliability across global deployments. You will collaborate with cross-functional teams to deliver cutting-edge infrastructure for AI and high-performance computing workloads.

Requirements

  • An experienced systems engineer with a strong background in HPC, AI infrastructure, and data center engineering.
  • Deep technical knowledge of data center rack layout, compute and networking components.
  • A strategic mindset for system-level design.
  • Ability to collaborate across diverse technical domains.
  • Thrive in fast-paced environments.
  • Passionate about building efficient, scalable, and reliable compute platforms.
  • Bachelor’s or Master’s degree in Electrical Eng

Nice To Haves

  • Extensive experience in HPC, AI infrastructure, or data center systems engineering
  • Experience with liquid cooling or advanced thermal management
  • Experience with rack level power distribution
  • Contributions to open-source HPC or AI infrastructure projects

Responsibilities

  • Design scalable AI/HPC data center rack layouts including compute, storage, networking, power delivery and cooling
  • Translate high level requirements and architecture input into detailed rack designs
  • Design leading-edge thermal and power delivery for high-density deployments
  • Design intra-rack and rack-to-rack network connections to maximize overall cluster performance
  • Translate network architecture requirements and input into rack-level designs
  • Understand the network performance needs of different types of workloads
  • Understand advantages and performance trade-offs of network topologies for AI/HPC clusters
  • Design rack-level and data center power delivery infrastructure
  • Define power budgets, redundancy schemes, and fault tolerance mechanisms
  • Understand differences in power delivery and regulatory requirements in global locations, e.g. U.S., EMEA, Asia and other countries
  • Translate storage architecture requirements and input into rack-level designs
  • Design and optimize storage solutions to maximize AI/HPC cluster performance
  • Understand advantages and performance trade-offs of cluster storage solutions, e.g. Lustre, Ceph, etc.
  • Work across multiple organizations with subject matter experts from hardware, software, network, data center, and operations teams to deliver scalable, efficient, and reliable compute infrastructure.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service