This role is for one of our clients. Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities. In this role, you will review complex, multi-step evaluation tasks, validate their correctness, identify edge cases, and strengthen testing methodologies before benchmarks are deployed. You'll collaborate closely with AI researchers and task authors to improve evaluation quality, eliminate ambiguity, and ensure benchmark integrity. This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior