About The Position

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This is a project-based opportunity, not permanent employment. We are building a dataset to evaluate AI coding agents, assessing their performance on real-world developer tasks. You will create challenging tasks and evaluation criteria within realistic simulated environments. This involves building realistic developer environments with a codebase, infrastructure, and context (tickets, docs, conversations) that form a believable development history. You will design tasks from intermediate states of these environments, crafting prompts, defining what "solved" means, and ensuring tasks are solvable by an AI agent. Additionally, you will write tests to verify agent solutions, ensuring they accept all valid approaches and reject incorrect ones. You will iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining until the evaluation is fair and robust. This role is not data labeling, prompt engineering, or writing code from scratch, as the AI agent will write most of the code; your role is to guide and evaluate.

Requirements

  • 8+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Responsibilities

  • Build realistic developer environments, including a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history.
  • Design tasks from intermediate states of these environments, crafting prompts, defining what "solved" means, and ensuring the task is solvable by an AI agent.
  • Write tests that verify agent solutions, accepting all valid approaches and rejecting incorrect ones.
  • Iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining until the evaluation is fair and robust.

Benefits

  • Up to $200/hr equivalent compensation
  • Flexible schedule
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service