This role is for one of our clients. Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Software Engineers to design and validate sophisticated engineering tasks that mirror real-world software development challenges. In this role, you will create complex, multi-step software engineering benchmarks that require implementation, debugging, environment configuration, and technical reasoning. Working alongside AI researchers, you'll help identify where advanced AI coding systems succeed, where they fail, and how benchmark quality can be continuously improved. This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level