GPU Infrastructure Engineer

RuneMountain View, CA
$175,000 - $260,000Onsite

About The Position

Every solar and wind power plant is a latent data center. Rune connects wasted power to the compute that needs it. Solar plants generate electricity the grid can't take: clipped by inverters, curtailed by operators, gone before it ever reaches a meter. Meanwhile, AI needs more power than new construction can deliver in time. Rune deploys hardware, software, and financial solutions to put this power to work. We build modular micro data centers that install directly at solar sites, bypassing the traditional grid entirely and turning stranded power into compute. We're backed by Spark Capital, Union Square Ventures, and Lowercarbon Capital. We're looking for a hands-on engineer to own customer-facing technical support for our GPU clusters: networking, node provisioning, and GPU health. You'll be the primary responder to ensure our customers’ clusters are performing as expected, including diagnosing issues, resolving them, or routing them correctly, without needing to hand off the technical judgment to someone else.

Requirements

  • Hands-on experience with RDMA fabrics — RoCEv2 specifically, including ECN/DCQCN, PFC, and DCBX trust-mode configuration on Mellanox/NVIDIA ConnectX NICs.
  • Experience debugging multi-node GPU interconnect performance with NCCL and RDMA benchmarking tools (perftest, nccl-tests).
  • Working knowledge of GPU interconnect topology (NVLink/NVSwitch, NUMA/PCIe affinity) and how it affects distributed training performance.
  • Experience reading GPU health telemetry (NVIDIA DCGM, nvidia-smi, Xid codes) and managing driver/firmware compatibility across a fleet.
  • Strong Linux systems administration, including SSH/host security tooling (fail2ban, sshd) and fleet automation (Ansible or equivalent).
  • Comfortable owning live, customer-facing escalations on production systems, with clear written and verbal communication under time pressure.

Nice To Haves

  • Familiarity with native InfiniBand fabrics (OpenSM/UFM) in addition to RoCEv2.
  • Experience with scheduler-level network integration (Slurm, Kubernetes, Run:AI), including SR-IOV and multi-NIC bonding.
  • Prior experience in a vendor-facing escalation role (e.g., NVIDIA/Mellanox support) or as a Field Applications Engineer

Responsibilities

  • Serve as the primary technical responder for customer-reported networking, GPU-interconnect, and node issues on live clusters.
  • Diagnose multi-node NCCL and RDMA performance regressions, including issues sourced to GPUDirect RDMA contention, and separate confirmed findings from working hypotheses when talking to customers.
  • Configure and validate RoCEv2 congestion control (ECN/DCQCN, PFC, DCBX) and reconcile host-side NIC settings against the switch fabric's actual behavior.
  • Diagnose GPU hardware faults using DCGM telemetry and Xid error codes, distinguishing genuine hardware failures (RMA) from driver, firmware, or configuration issues.
  • Own driver and firmware version compatibility across the fleet: GPU driver, NIC firmware (OFED/MOFED), kernel, and validate compatibility before and after updates.
  • Provision, re-provision, and decommission GPU nodes: power-cycling, secure wipe, re-imaging, BIOS/NUMA/PCIe topology validation, and acceptance testing before handoff to a customer.
  • Triage edge-network and jumpbox access issues affecting customer provisioning workflows (e.g., automated blocking of legitimate high-concurrency SSH/Ansible traffic).
  • Write clear incident updates and post-resolution summaries for customer engineering teams, and feed recurring issues back into provisioning defaults and acceptance-test procedures.

Benefits

  • Competitive base and strategic ownership in what we believe will be one of the largest infrastructure build-outs of the decade.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service