This role is for one of our clients. Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking researchers from computational STEM disciplines—as well as computationally intensive social sciences and humanities—to bring the rigor of real-world research into AI evaluation. In this role, you will transform scientific methodologies such as experimental design, hypothesis testing, and data-driven analysis into sophisticated, multi-step benchmark tasks that challenge state-of-the-art AI systems. Working closely with AI researchers, you'll help uncover subtle reasoning errors and methodological flaws that only experienced researchers can identify. This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior