Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops

Advanced Micro Devices, IncSanta Clara, CA

About The Position

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. AMD is looking for an influential software engineer who is passionate about improving the performance of key applications and benchmarks. You will be a member of a core team of incredibly talented industry specialists and will work with the very latest hardware and software technology.

Requirements

  • Exceptional technical expertise, bridging deep proficiency in high-performance C++ software engineering and low-level GPU programming with a robust understanding of Large Language Models (LLMs) and AI systems.
  • Ability to bridge kernel engineering with AI post-training (RL) experience.
  • Mastery in designing complex, scalable systems using modern C++.
  • Fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and kernel optimization to maximize hardware performance.
  • Significant hands-on experience in large-scale C++/HIP/CUDA projects, such as contributing to the ROCm ecosystem (e.g., rpp, MIVisionX, rocAL, rocdecode, rocjpeg), CUDA libraries (e.g., CV-CUDA, cuDNN, NCCL), or the C++/HIP/CUDA core of ML frameworks like PyTorch, TensorFlow, or JAX.
  • Deep understanding of LLMs, including transformer architectures, attention mechanisms, and the full model lifecycle.
  • Hands-on experience in advanced model alignment and post-training techniques like Supervised Fine-Tuning (SFT) and Reinforcement Learning (e.g., RLHF, GRPO).
  • Familiarity with cutting-edge LLM trends such as Mixture-of-Experts (MoE) architectures, inference optimizations (e.g., quantization, speculative decoding), and modern application patterns like Agentic AI systems (e.g. AlphaEvolve for code/kernel generation).
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.

Nice To Haves

  • Experience and interest in code generation and/or self-improving LLMs.
  • Lengthy professional software development experience in performance-critical environments.
  • Extensive hands-on experience in GPU programming (HIP/CUDA) and optimizing deep learning kernels and operators.
  • Computer vision expertise.
  • A fundamental understanding of GPU architecture and memory hierarchy, used to diagnose and resolve complex performance bottlenecks.
  • Expert-level proficiency in modern C++ and object-oriented design.
  • Deep experience using GPU profiling and performance analysis tools (e.g., AMD ROCm Profiler, NVIDIA Nsight) to diagnose and resolve complex bottlenecks in distributed, multi-GPU systems.
  • Deep knowledge of transformer architectures, attention mechanisms, and modern AI systems (Generative AI, Agentic AI).
  • Hands-on experience optimizing the post-training and inference pipelines of Large Language Models (LLMs).
  • Strong technical ownership, communication, and problem-solving skills with a track record of delivering complex technical solutions.
  • Experience or deep expertise with the AMD ROCm/HIP ecosystem.
  • Relevant publications in AI/ML, GPU computing, or system optimization.

Responsibilities

  • Architect and Drive the AI Software Stack: Establish best practices and optimize performance from the lowest-level GPU kernels to large-scale distributed systems, shaping the foundational software for AMD hardware. Leverage cutting-edge Large Language Models (LLMs) and agent-based technologies to accelerate the development and performance enhancement of the AMD ROCm ecosystem.
  • Accelerate Foundational Models: Directly accelerate cutting-edge applications like foundation models (LLMs) and autonomous AI agents, ensuring AMD is the platform of choice for the most demanding workloads.
  • Innovate Across Hardware and Software: Contribute to the entire co-design lifecycle, from influencing future GPU architectures to developing groundbreaking software for new accelerators and collaborating with the broader AI community.
  • Mentor others and effectively communicate ideas to shape the future of AI at AMD.

Benefits

  • AMD benefits at a glance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service