You'll own the health, reliability, and performance of Evergrid's AMD-only GPU compute clusters. You're the primary custodian of our high-density accelerator environments. The work spans hardware operations, Linux systems engineering, distributed infrastructure, and ML workloads. It covers GPU bring-up and kernel-level debugging, as well as maintaining and optimizing the ROCm-based ML stack behind production-scale AI. If you like getting maximum performance out of hardware, debugging GPUs at scale, and shipping world-class AI infrastructure, this role is for you.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior