We are seeking an AI Research Infrastructure Engineer to operate, scale, and continuously improve the shared GPU and HPC compute platform behind our AI, ML, and HPC research. You will own the day-to-day health of our SLURM and GPU clusters and work hands-on with researchers to get demanding workloads—including large-scale multi-GPU and multi-node training—running reliably and efficiently. This is a research-enablement role, not a traditional systems-administration role. It requires research literacy—a working understanding of how modern models are trained and where they bottleneck—so you can partner with researchers as a technical peer and directly accelerate their work. The focus is on enabling and operating research infrastructure, not pursuing an independent research agenda. Your scope spans both internal and externally visible compute clusters across the research organization, including AMD's university program clusters and interfacing with external academic collaborators. You may also help coordinate the contractors and systems administrators supporting the environment. Familiarity with modern agentic engineering workflows—tools such as Claude Code, Codex, or Cursor—is also expected.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior