Senior HPC & AMD Infrastructure Engineer

Evergrid•New York, NY
•$180,000 - $220,000•Hybrid

About The Position

You'll own the health, reliability, and performance of Evergrid's AMD-only GPU compute clusters. You're the primary custodian of our high-density accelerator environments. The work spans hardware operations, Linux systems engineering, distributed infrastructure, and ML workloads. It covers GPU bring-up and kernel-level debugging, as well as maintaining and optimizing the ROCm-based ML stack behind production-scale AI. If you like getting maximum performance out of hardware, debugging GPUs at scale, and shipping world-class AI infrastructure, this role is for you.

Requirements

  • 5+ years in HPC, GPU cluster operations, Linux systems engineering, or similar roles.
  • A bachelor's or master's in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Deep hands-on experience with AMD MI-series GPUs, including driver and kernel-level debugging.
  • Strong grasp of Linux internals, kernel modules, hardware bring-up, and performance tuning.
  • Experience securing and operating production infrastructure: VPNs, firewalls, SSH, and identity systems.
  • Proficiency in Bash and Python for automation, tooling, and operations.
  • Strong familiarity with ML stacks and runtime behavior in ROCm environments (ROCm/HIP, MIOpen, RCCL, PyTorch, JAX).
  • Experience debugging high-performance networking and RDMA (InfiniBand or RoCE), including cluster-level communication failures that affect distributed training.

Nice To Haves

  • Schedulers and orchestration (Slurm, Kubernetes).
  • Model serving and inference optimization on ROCm (vLLM, SGLang).
  • Configuration management and IaC (Ansible, SaltStack, Terraform).
  • Supporting ML research or production AI teams at a startup or high-growth company.

Responsibilities

  • System health and reliability (SRE)
  • Primary on-call response for outages, GPU failures, node crashes, and cluster-wide incidents.
  • Being the key point person for POC and active customers' GPU clusters.
  • Fast diagnosis and resolution that minimizes downtime and keeps SLA-level reliability.
  • Monitoring for GPU health, thermals, PCIe topology, memory errors, and cluster load.
  • Repairs, RMAs, and physical maintenance, coordinated with data center operators, hardware vendors, and on-site technicians.
  • Installing, patching, and maintaining Linux (Ubuntu, CentOS, RHEL) across large GPU node fleets.
  • Kernel tuning, consistent OS configuration, and fleet automation at scale.
  • Secure networking: VPNs, firewalls (iptables/firewalld), SSH hardening, and routing.
  • Identity and access systems (LDAP, FreeIPA, Active Directory).
  • Distributed storage (NFS, GPFS, Lustre).
  • Deployment and bring-up of new GPU nodes, including BIOS configuration, NUMA tuning, and topology validation.
  • AMD GPU drivers, kernel modules, and the ROCm runtime across production fleets.
  • The AMD ML stack: ROCm, PyTorch (ROCm builds), JAX (ROCm/XLA), RCCL, hipBLAS/hipDNN, MIOpen, and supporting runtimes.
  • Debugging complex failures across GPUs, compilers, ML frameworks, and distributed training and inference. Examples include RCCL hangs, HIP memory faults, ROCm kernel crashes, framework build and link issues, and vLLM build failures on ROCm.
  • Infrastructure that supports both research iteration and production reliability, built with the ML and platform teams.

Benefits

  • Health insurance
  • Paid time off and paid holidays
  • Home office stipend
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service