About The Position

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This is a project-based opportunity, not permanent employment. The role involves creating challenging tasks and evaluation criteria within realistic simulated developer environments. This includes building virtual companies with codebases, infrastructure, and context, designing tasks from intermediate states of these environments, writing tests to verify AI agent solutions, and iterating on tasks and tests based on QA feedback. This role is not data labeling, prompt engineering, or writing code from scratch; the AI agent writes most of the code, and the engineer guides and evaluates.

Requirements

  • 8+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Responsibilities

  • Build realistic developer environments (virtual company with codebase, infrastructure, context).
  • Design tasks from intermediate states of these environments, including crafting prompts and defining 'solved' criteria.
  • Write tests that verify AI agent solutions, accepting all valid approaches and rejecting incorrect ones.
  • Iterate on tasks and tests based on QA feedback, analyzing failures, and refining evaluations for fairness and robustness.

Benefits

  • Paid per accepted task
  • Rate depends on qualification tier and efficiency
  • Up to the equivalent of $200/hr
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service