Thinking Machines is seeking a network engineer to manage the foundational network layers critical for large-scale AI training and inference. This role involves ensuring the reliability of interconnects across extensive GPU fabrics, including RDMA/RoCE between nodes and NVLink/NVSwitch within nodes. The position requires hands-on, cross-stack debugging, building instrumentation and tooling to enhance debugging efficiency, and acting as a technical liaison with cloud providers' networking teams to resolve issues. The ultimate goal is to provide researchers with a dependable infrastructure. This is an evergreen role, meaning applications are continuously reviewed for current and future opportunities. While an immediate role may not always be available, interested candidates are encouraged to apply. Reapplication is permitted every six months, and separate postings for specific project needs may also arise.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Associate degree