2027 PhD AI Training Systems and Performance Engineer Intern/Co-Op

Advanced Micro Devices, Inc•San Jose, CA
•Hybrid

About The Position

As an AMD intern and co-op, you’ll be placed at the epicenter of the AI ecosystem, working alongside experts and industry pioneers. You’ll do important work, learn new skills, expand your network, and gain real-world experience on projects that impact millions of end-users worldwide. Whether you’re an undergrad or a PhD student, your contributions matter—and your experience here will be a launchpad for what comes next. We are seeking a highly motivated AI Training Systems & Performance Engineering PhD Intern/Co-Op to join our team. In this role, you will help accelerate the adoption and optimization of cutting-edge AI training workloads on AMD Instinct™ GPUs while contributing to the next generation of AI software performance solutions. You will work alongside software engineers, architects, and AI specialists to bring up new training workloads, analyze performance bottlenecks, and develop innovative tooling that improves scalability, efficiency, and developer productivity. We will involve you in developing and optimizing large-scale AI training and fine-tuning workloads running on AMD GPU platforms. You will help bring up newly released foundation models and training frameworks, adapting implementations and training recipes for AMD hardware while establishing reproducible correctness and performance baselines. We will work with you to profile and analyze AI workloads, identifying bottlenecks across GPUs, CPUs, memory systems, networking, and communication infrastructure. You will develop tools and workflows that automate training setup, debugging, performance analysis, and optimization using LLM-powered agents and agentic AI techniques. You will investigate and implement optimization strategies that improve training throughput, GPU utilization, memory efficiency, and scalability across distributed multi-GPU environments. We will expose you to advanced profiling and performance analysis tools to benchmark and tune AI frameworks, libraries, SDKs, and applications running on AMD platforms. You will collaborate with software engineers and architects to evaluate emerging AI models, distributed training techniques, and performance optimization opportunities. Your work will help transform successful experiments into reusable workflows, best practices, and software improvements that enhance out-of-the-box AI training performance on AMD hardware.

Requirements

  • Currently pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, Machine Learning, or a related technical discipline.
  • Strong programming experience in Python and/or C++.
  • Hands-on experience implementing, training, and debugging deep learning models using frameworks such as PyTorch, JAX, TensorFlow, vLLM, or SGLang.
  • Experience with one or more of the following areas: Distributed training systems, Data, tensor, pipeline, expert, or context parallelism, GPU performance optimization, AI systems software, High-performance computing (HPC), Large language model training and fine-tuning, Agentic AI or LLM-powered automation.
  • Understanding of transformer-based architectures, mixture-of-experts models, and modern LLM training techniques.
  • Experience profiling workloads using performance analysis tools such as PyTorch Profiler, ROCm Profiler, VTune, Nsight, or similar tools.
  • Familiarity with distributed training technologies and communication libraries such as MPI, NCCL/RCCL, OpenMP, or related frameworks.
  • Understanding of GPU architecture, memory systems, communication bottlenecks, and performance tuning methodologies.

Nice To Haves

  • Experience identifying and resolving compute, memory, data-loading, or communication bottlenecks in large-scale AI workloads is preferred.
  • Experience with ROCm, HIP, Triton, GPU kernel optimization, or AI systems software development is a plus.
  • Publications in AI, Machine Learning, High Performance Computing, Computer Architecture, or related research areas are a plus.

Responsibilities

  • Develop and optimize large-scale AI training and fine-tuning workloads running on AMD GPU platforms.
  • Bring up newly released foundation models and training frameworks, adapting implementations and training recipes for AMD hardware while establishing reproducible correctness and performance baselines.
  • Profile and analyze AI workloads, identifying bottlenecks across GPUs, CPUs, memory systems, networking, and communication infrastructure.
  • Develop tools and workflows that automate training setup, debugging, performance analysis, and optimization using LLM-powered agents and agentic AI techniques.
  • Investigate and implement optimization strategies that improve training throughput, GPU utilization, memory efficiency, and scalability across distributed multi-GPU environments.
  • Benchmark and tune AI frameworks, libraries, SDKs, and applications running on AMD platforms using advanced profiling and performance analysis tools.
  • Collaborate with software engineers and architects to evaluate emerging AI models, distributed training techniques, and performance optimization opportunities.
  • Transform successful experiments into reusable workflows, best practices, and software improvements that enhance out-of-the-box AI training performance on AMD hardware.

Benefits

  • AMD benefits at a glance

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume

What This Job Offers

Job Type

Full-time

Career Level

Intern

Education Level

Ph.D. or professional degree

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service