This role is for one of our clients. Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Data Scientists and Quantitative Analysts to bring real-world analytical rigor to AI evaluation by designing sophisticated benchmark tasks based on practical data science workflows. In this role, you will create complex, multi-step analytical challenges that mirror real research and business scenarios—from cleaning datasets and comparing statistical methods to interpreting results and presenting actionable insights. Working closely with AI researchers, you'll help identify where advanced AI models succeed, where they fail, and how evaluation benchmarks can better measure analytical reasoning. This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level