Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This is a project-based opportunity, not permanent employment. We are building a dataset to evaluate AI coding agents, assessing their performance on real-world developer tasks. You will create challenging tasks and evaluation criteria within realistic simulated environments. This involves building realistic developer environments with a codebase, infrastructure, and context (tickets, docs, conversations) that form a believable development history. You will design tasks from intermediate states of these environments, crafting prompts, defining what "solved" means, and ensuring tasks are solvable by an AI agent. Additionally, you will write tests to verify agent solutions, ensuring they accept all valid approaches and reject incorrect ones. You will iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining until the evaluation is fair and robust. This role is not data labeling, prompt engineering, or writing code from scratch, as the AI agent will write most of the code; your role is to guide and evaluate.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Senior
Education Level
No Education Listed