Training a frontier model is as much an infrastructure problem as it is a modeling problem. When a run fails at hour 47 of a 72-hour job, the question is almost never "what's wrong with the code" — it's "what happened on the cluster." Node failures, network timeouts, storage bottlenecks, and scheduling decisions determine whether training runs succeed or fail, how much they cost, and how fast teams can iterate. Until now, the developer tools and the infrastructure telemetry have lived in completely separate worlds. The CoreWeave acquisition changes that. W&B can now build products that connect application-level experiment tracking with cluster health, hardware telemetry, and compute orchestration — a surface area that's hard to assemble without owning both the developer tools and the infrastructure. Reporting to the Director of Product Management, this Staff Product Manager will own the strategy and execution for the products at this intersection: W&B Launch (job submission and compute orchestration), infrastructure insights that leverage CoreWeave's unique telemetry and platform capabilities, and new product concepts — including agentic and automated research workflows — that only exist because W&B and CoreWeave are now one company. Sandboxes (managed development environments) are also part of this surface, but the core opportunity is building infrastructure-native developer tools that take advantage of the combined platform. You'll be working directly with frontier model builders — teams training on thousands of GPUs — to understand what they need and ship it. If you're energized by solving infrastructure challenges that directly accelerate deep learning workflows, collaborating with the world’s best AI teams, and delivering tools that move the industry frontier forward, his role is for you.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed