We're looking for a Research Scientist to set and pursue a research agenda for reliable long-horizon agents operating inside real enterprises. Frontier labs optimize for general capability, and the public agent benchmarks are mostly sandboxes. Very little rigorous work exists on what it takes for an agent to reason across a system with nineteen years of undocumented decisions in it, plan a change across forty coupled steps, recover when step twelve reveals the model of the world was wrong, and be right often enough that a CFO signs the go-live. Almost nobody has the landscapes, the traces, or the customers to study it. We do. Two properties make this an unusually good research setting. First, much of the task space is verifiable — a transformation either produces a system that builds, passes regression, and behaves equivalently, or it doesn't. That's a real reward signal, not a preference model. Second, the parts that aren't verifiable are where the interesting work is: is this reconciliation correct, or merely plausible? Was retiring that capability the right call? Designing reward and evaluation across that boundary is the central research question here. You'll invent methods rather than only apply them, work with Research Engineers who help you run at scale, and hear from a product team within weeks whether you were right. We'd like you to publish. Not everything, and never at the expense of shipping — but the work here is novel enough to be worth writing down. One thing worth knowing up front: we post-train open-weight models on rented clusters and buy more compute when a result justifies it. We're constrained relative to a frontier lab. If your research only works at ten thousand GPUs, this is the wrong place.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior