AI Research Infrastructure Engineer

Advanced Micro Devices, IncAustin, TX
Onsite

About The Position

We are seeking an AI Research Infrastructure Engineer to operate, scale, and continuously improve the shared GPU and HPC compute platform behind our AI, ML, and HPC research. You will own the day-to-day health of our SLURM and GPU clusters and work hands-on with researchers to get demanding workloads—including large-scale multi-GPU and multi-node training—running reliably and efficiently. This is a research-enablement role, not a traditional systems-administration role. It requires research literacy—a working understanding of how modern models are trained and where they bottleneck—so you can partner with researchers as a technical peer and directly accelerate their work. The focus is on enabling and operating research infrastructure, not pursuing an independent research agenda. Your scope spans both internal and externally visible compute clusters across the research organization, including AMD's university program clusters and interfacing with external academic collaborators. You may also help coordinate the contractors and systems administrators supporting the environment. Familiarity with modern agentic engineering workflows—tools such as Claude Code, Codex, or Cursor—is also expected.

Requirements

  • Operations- and infrastructure-minded, with a strong bias toward reliability, automation, usability, and reducing operational overhead for researchers.
  • Strong hands-on Linux systems administration in shared/multi-user, production compute environments.
  • Practical, hands-on experience with the SLURM workload manager—partitions, scheduling policy, accounting, node management, and tuning.
  • Experience operating GPU, HPC, or AI/ML research compute environments.
  • Research literacy: a working understanding of how large models are trained and where they bottleneck, alongside researchers as a technical peer on multi-GPU/multi-node workloads.
  • Experience with Docker/containers and Kubernetes, plus automation and system health monitoring.
  • Familiarity with agentic engineering workflows using tools such as Claude Code, Codex, or Cursor.

Nice To Haves

  • DevOps practices: GitHub Actions, self-hosted runners, CI/CD pipelines, and infrastructure automation / Infrastructure as Code.
  • Container registries and reproducible runtime environments.
  • Shared storage, networking, and high-speed interconnects (InfiniBand, RoCE) in multi-user clusters.
  • AMD GPU platforms and the ROCm/RCCL stack—hardware bring-up, architecture-aware debugging, validation, or performance workflows.
  • Supporting developer- or researcher-facing platforms.

Responsibilities

  • Own day-to-day operations of the SLURM-managed GPU and HPC clusters, ensuring high availability, utilization, and performance across a multi-user research environment.
  • Partner directly with researchers as a technical peer to get demanding multi-GPU and multi-node workloads running and optimized.
  • Operate and improve the broader compute platform—shared storage, networking, containers, and monitoring—and build automation and self-service workflows that reduce friction for researchers.
  • Support GPU platform and hardware bring-up: validation, enablement, debugging, and operational readiness.
  • Manage AMD's university program clusters and support external academic collaborators alongside internal research users, spanning both internal and externally visible compute clusters.
  • Lead incident response and root-cause analysis, and build runbooks and preventive practices to improve reliability.
  • Set operational standards and best practices, and help coordinate the contractors and systems administrators supporting the environment.
  • Apply agentic coding and operations workflows to improve velocity across deployment, troubleshooting, documentation, and infrastructure management.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service