ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: We are seeking a highly motivated and skilled Principal AI Cluster Performance Validation Engineer to join our dynamic team. In this role, you will be at the forefront of optimizing and achieving peak performance for GPU clusters. The focus of this role is the RDMA networks used in AI Clusters, understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures, parallel clustered computing, and hands-on experience in system level performance tuning and debug methodologies. THE PERSON: The team fosters and encourages continuous technical innovation to showcase successes as well as facilitate continuous career development. A seasoned professional who enjoys hands-on problem-solving. In this role, you’ll shape long-term strategy, drive feature enablement and jump in to tackle challenges head-on. You’ll have a direct impact on performance, automation, and validation, while staying ahead of industry trends to provide strategic insights to senior management. The person should be experienced in debugging complex HW/FW and clustered configurations.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal