Handshake AI works directly with frontier labs on their most consequential data, evaluation, and post-training challenges, building the systems that turn expert human knowledge into the data and evaluations that make frontier models better. This role is for someone who is deeply passionate about building software with agents and creating the evaluations that reveal where those agents truly succeed and fail. You will design and publish coding benchmarks that meaningfully challenge state-of-the-art agents. You will work with AI researchers, software engineers, and domain experts to turn difficult, real-world software tasks into high-signal evaluation environments, datasets, verifiers, and feedback systems. The work will help shape both how the frontier evaluates coding agents today and where the field goes next. Early members of the team will have unusual influence over our technical direction, operating culture, and the open-source benchmarks, software, and research products we build. We care more about demonstrated technical depth, judgment, and a builder's mindset than a specific title, degree, or career path.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed