Triton Compiler and Kernel Software Engineer

Advanced Micro Devices, Inc•San Jose, CA
•Hybrid

About The Position

We are seeking a Senior Triton Compiler and Kernel Engineer to advance Triton performance and capabilities on AMD GPUs. Triton is an open-source language and compiler for developing high-performance GPU kernels in Python. It is a critical layer in the AI software stack, connecting frameworks and workloads to GPU hardware. Triton is strategic to AMD’s AI roadmap, and AMD is investing fully in making it a first-class platform for current and future AMD GPUs. You will work across GPU architecture, compilers, kernels, multi-GPU communication, and AI frameworks while contributing to upstream Triton and AMD’s ROCm software stack.

Requirements

  • Strong experience in GPU architecture and programming
  • Strong experience in Compiler development and optimization
  • Strong experience in High-performance GPU kernels
  • Strong experience in Multi-GPU communication and collective operations
  • Strong experience in Distributed AI training and inference
  • Strong experience in AI workload and framework performance
  • Strong experience in Low-level performance analysis
  • Ability to reason across the stack—from distributed AI algorithms and Triton programs to compiler transformations, communication libraries, generated instructions, and GPU hardware.

Nice To Haves

  • Experience with AMD GPU architecture and ROCm is highly desirable.
  • Experience optimizing kernels with Triton, HIP, CUDA, or GPU assembly.
  • Experience with Triton, LLVM, MLIR, or another optimizing compiler.
  • Knowledge of GPU execution models, memory hierarchies, synchronization, and instruction pipelines.
  • Experience with collective communication, distributed programming, and libraries such as RCCL or NCCL.
  • Understanding of communication topologies, interconnects, synchronization, and communication–computation overlap.
  • Familiarity with distributed training, inference, tensor parallelism, expert parallelism, or pipeline parallelism.
  • Experience using AI coding tools and autonomous agents to accelerate software development, debugging, benchmarking, and kernel performance tuning.
  • Familiarity with AI primitives, reduced-precision formats, and performance profiling.
  • Contributions to complex or open-source software projects.
  • Strong analytical, debugging, communication, and collaboration skills.

Responsibilities

  • Develop and optimize Triton compiler support for AMD GPUs.
  • Improve compiler lowering, optimization, scheduling, and code generation.
  • Create high-performance kernels for attention, GEMM, MoE, and other AI workloads.
  • Develop and optimize multi-GPU kernels and communication primitives.
  • Enable new AMD GPU and interconnect capabilities through effective Triton abstractions.
  • Analyze compute, memory, communication, occupancy, register usage, and generated code.
  • Resolve complex correctness and performance issues across single- and multi-GPU workloads.
  • Collaborate with GPU architecture, ROCm, PyTorch, and AI framework teams.
  • Contribute designs and implementations to upstream Triton and LLVM/MLIR.
  • Provide technical leadership, code reviews, and mentoring.

Benefits

  • AMD benefits at a glance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service