The mission of Thinking Machines is to build AI that extends human will and judgment. Evaluation is one of the most important pillars of building frontier AI systems. It guides research direction, powers experimentation, and helps us understand whether changes to data and training are improving the capabilities and behaviors we care about. To support this work, researchers need a powerful, self-serve platform that makes it easy to author evaluations, run them or reproduce them reliably at scale, and extract insight from the results. The platform must support both standardized external benchmarks and fast-moving internal evaluations, many kinds of tasks and graders, and inspection from aggregate metrics down to individual model trajectories. In this role, you will design and build this platform end to end. You will work across Python frameworks, data pipelines, APIs, and user-facing applications, and collaborate closely with pre-training, post-training, and applied teams to improve how we evaluate models and turn results into research decisions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level