We are hiring senior full-stack engineers to help build and refine next-generation coding environments that evaluate AI agents. This work directly supports the development of robust grading systems that accurately measure how effectively an AI agent solves complex coding tasks. The goal is to create secure evaluation suites that cannot be bypassed or easily cheated by the models. Throughout this long-term engagement, you will work remotely with real website clones to identify bugs and write technical specifications. You will build comprehensive automated test suites designed to grade AI-generated code. This involves analyzing how an agent interacts with a given task and ensuring the grading logic is watertight. The process requires deep technical scrutiny and creative problem-solving to account for unpredictable AI behavior.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Senior
Education Level
No Education Listed