We are building AI to simulate the world through merging art and science. World models offer the most clear path to general-purpose simulation, changing how stories are told, how scientific progress is made and how the next frontiers of humanity are reached. Our team consists of creative, open minded, caring and ambitious people who are determined to change the world. We aspire to continuously build impossible things and our ability to do so relies on building an incredible team. If you are driven to do the same, we'd love to hear from you. We're looking for an ML infrastructure engineer to own model evaluation at Runway, end to end. Every decision we make about a model – which checkpoint to keep training, what to ship to millions of users, which datamix and architecture shows the most promise – rests on evals. Today that work is spread across the organization. You'll turn it into one platform. Your job will be to design and build the systems that generate samples at scale, score them with automated metrics and human annotations, track results across checkpoints and releases, and put the answers in front of researchers in minutes rather than days. You'll define what "better" means operationally – how we measure it, how confident we are, and how a result becomes a ship/no-ship decision. You'll be embedded within research teams as a member of ML Platform, and you'll set the technical direction for evals across the company. This is an extremely high-leverage role: the quality of our models is bounded by how well we can measure them.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed