We're hiring a Software Engineer to sit at the day-to-day interface between research and infrastructure. This is a generalist role with broad scope: you'll be one of the people always in the room for infra decisions on the ML systems side, with a clear enough view of upcoming compute needs to see support burden coming before it arrives. You'll also be one of the people who babysits hero runs — the ones at 2am when a 4k-GPU job hits a weird Xid and someone needs to decide, quickly and correctly, whether to drain the node, restart the job, or escalate to NVIDIA before the run loses checkpoints. That kind of judgment, built across the kernel, the network, the scheduler, and the application layer, is the core of the job.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed