Senior Network Engineer – GPU Cluster Networking

Advanced Micro Devices, Inc•San Jose, CA
•Hybrid

About The Position

Join AMD's IT Systems Engineering team and help build the networking foundation powering some of the world's most advanced AI and high-performance computing environments. In this role, you will design, automate, scale, and operate high-performance networks supporting AMD Instinct GPU clusters used for AI training, inference, and HPC workloads. Working closely with AI, platform, systems, and data center engineering teams, you will influence next-generation network architecture, optimize end-to-end infrastructure performance, and help shape the future of accelerated computing at AMD. This role also supports AMD's global campus networking environment spanning locations around the world.

Requirements

  • Bachelor's or Master's degree in Computer Engineering, Computer Science, Information Technology, or a related field, or equivalent practical experience.
  • CCIE or comparable advanced networking certifications are valued.

Nice To Haves

  • Experience designing and operating large-scale data center networks supporting AI, GPU, HPC, cloud, or distributed computing environments.
  • Strong understanding of modern networking technologies including Ethernet fabrics, BGP, ECMP, QoS, EVPN, VXLAN, and related routing and switching protocols.
  • Experience with RoCEv2, RDMA, network performance optimization, observability, monitoring, and troubleshooting.
  • Experience with network automation, scripting, infrastructure-as-code, or operational tooling.
  • Experience with enterprise and data center networking platforms such as Juniper, Cisco, Arista, Aruba, ClearPass, Aruba Central, Silver Peak, Infoblox, or similar technologies.
  • Experience collaborating across infrastructure, platform, systems, and data center engineering teams.

Responsibilities

  • Design, deploy, operate, and continuously improve data center networks supporting large-scale AI and HPC environments.
  • Develop scalable network architectures and optimize performance across compute, storage, and networking infrastructure.
  • Automate network operations and implement solutions that improve reliability, efficiency, and scalability.
  • Lead troubleshooting, incident response, root cause analysis, and continuous improvement initiatives.
  • Plan and execute network expansions, technology upgrades, and infrastructure migrations.
  • Support and enhance global campus networking technologies including wireless, LAN, WAN, and SD-WAN environments.
  • Partner with cross-functional engineering and operations teams to deliver highly available network solutions.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service