HPC Systems Engineer - AI Workloads

Advanced Micro Devices, IncSan Jose, CA
Onsite

About The Position

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: We are seeking an AI Systems Engineer to join our AMD IT compute platforms engineering team. The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers.

Requirements

  • Significant experience in working across a globally distributed organization.
  • Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and Cluster interconnect with 400G networking.
  • Demonstrated experience with AI workload schedulers and allocation optimization.
  • Strong organizational, problem-solving, and troubleshooting skills, with the ability to manage multiple projects simultaneously.
  • Excellent verbal and written communication skills, with the ability to collaborate effectively with team members and stakeholders at all levels of the organization.

Nice To Haves

  • Experience in developing Python based AI apps and UI
  • HPC infrastructure engineering for AI/HPC domain
  • SLURM and Kubernetes management
  • Managing GPU clusters optimizing GPU-based services/tools/software
  • Experience in creating web services with HPC backend (like AI)
  • Automation/monitoring tool - Ansible / Saltstack, Terraform, Prometheus, Grafana

Responsibilities

  • Develop, implement, and maintain GPU-based clusters, ensuring optimal performance
  • Administer ML/AI platforms – Distributed ML services, LLMs and AI inferencing, by managing deployments, resource allocation, monitoring, and security.
  • Automate system provisioning and Cluster management end to end
  • Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technical expertise.
  • Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards.
  • Use AI/ML to continuously improve internal processes and tools that are used in end-to-end delivery of your services in this team

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service