Sr. Data Center GPU Validation and Debug Engineer

Advanced Micro Devices, IncAustin, TX
Hybrid

About The Position

We are seeking an experienced engineer to validate and debug data center GPU products across the complete software and hardware stack, spanning workload behavior and performance qualification as well as platform health and root-cause isolation. The role demands strong technical judgment, disciplined failure isolation, and the ability to trace unexpected behavior from application to hardware, working across firmware, kernel, compiler, library, and framework teams to drive issues to resolution.

Requirements

  • Strong Linux and systems-engineering experience, with C/C++, Python, and shell automation.
  • Experience with data center GPUs, accelerators, or comparable high-performance systems.
  • Understanding of GPU architecture, memory hierarchy, parallel execution, communication, and synchronization.
  • Experience debugging across software, firmware, kernel, and hardware boundaries.
  • Ability to analyze logs, traces, hardware telemetry, and performance counters.

Nice To Haves

  • Strong knowledge of Linux kernel drivers, PCIe, firmware interaction, and memory management.
  • Familiarity with GPU programming models and runtimes such as ROCm/HIP, RCCL, or CUDA equivalents.
  • Experience with AI training or inference frameworks such as PyTorch, vLLM, or SGLang.
  • Experience benchmarking AI inference, training, or HPC workloads, including roofline analysis and model-level profiling.
  • Experience with profiling tools, statistical analysis, and regression detection.
  • Familiarity with multi-GPU topology, PCIe/fabric interconnects, NUMA, GPU firmware, RAS, and rack-scale platforms.
  • Experience with system bring-up, qualification, or production data center operations.
  • Familiarity with CI systems, automated validation infrastructure, and fleet-scale testing.

Responsibilities

  • Debug GPU failures including hangs, crashes, memory faults, firmware errors, and performance regressions across the full software and hardware stack.
  • Reproduce failures and reduce them to minimal, actionable test cases.
  • Triage system-level interactions across GPU, CPU, memory, networking, power, and platform topology.
  • Validate AI, HPC, and communication workloads across GPU platforms, firmware, and software releases.
  • Characterize performance and identify compute, memory, and communication bottlenecks.
  • Analyze single- and multi-GPU behavior across workload configurations, topology, NUMA, power, and thermal conditions.
  • Establish benchmarks, baselines, and regression-detection methods; correlate microbenchmark results with real workload behavior.
  • Build automated diagnostics, stress tests, performance suites, and regression infrastructure.
  • Support customer-representative validation and escalation reproduction.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service