Network Engineer, Supercomputing

Thinking Machines LabSan Francisco, CA
$350,000 - $475,000Onsite

About The Position

Thinking Machines is seeking a network engineer to manage the foundational network layers critical for large-scale AI training and inference. This role involves ensuring the reliability of interconnects across extensive GPU fabrics, including RDMA/RoCE between nodes and NVLink/NVSwitch within nodes. The position requires hands-on, cross-stack debugging, building instrumentation and tooling to enhance debugging efficiency, and acting as a technical liaison with cloud providers' networking teams to resolve issues. The ultimate goal is to provide researchers with a dependable infrastructure. This is an evergreen role, meaning applications are continuously reviewed for current and future opportunities. While an immediate role may not always be available, interested candidates are encouraged to apply. Reapplication is permitted every six months, and separate postings for specific project needs may also arise.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Ability to thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where opportunities for improvement are identified.

Nice To Haves

  • Fluency with host-level debugging tools on Linux.
  • Strong communication skills, internally and with cloud providers.
  • Extensive experience with at least one of the following: Familiarity with cloud network primitives across at least two cloud providers.
  • Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
  • Statistical rigor in reliability reasoning — comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
  • A track record of writing tooling that made the next debugging session meaningfully faster.
  • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
  • Understanding of deep learning frameworks and their underlying system architectures.

Responsibilities

  • Reason about and validate GPU network fabric design across our deployments.
  • Debug RDMA / RoCEv2 across different NIC vendors, diagnosing collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
  • Own NVLink / NVSwitch interconnect, including fabric manager and IMEX health, link and lane errors, and their interaction with collectives.
  • Build host-level network instrumentation and use Linux tooling to create dashboards and alerts.
  • Navigate cross-cloud fabric quirks across providers and triage issues across NIC, driver, kernel, switch, and workload boundaries.
  • Drive escalations with cloud-provider networking teams, owning issues end-to-end until resolution.

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
  • Visa sponsorship
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service