About The Position

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This is a freelance, project-based opportunity, not permanent employment. The role involves creating challenging tasks and evaluation criteria within realistic simulated developer environments to assess the capabilities of AI coding agents. This includes building virtual companies with codebases, infrastructure, and development context, designing tasks from intermediate states, defining success criteria, and writing tests to verify AI agent solutions. The position requires iterating on tasks and tests based on QA feedback to ensure fair and robust evaluation.

Requirements

  • 8+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Nice To Haves

  • Deep understanding of where AI models fail in coding tasks.
  • Ability to create tasks that genuinely challenge advanced AI models.
  • Skill in writing tests that accept all correct solutions and reject incorrect ones.

Responsibilities

  • Build realistic developer environments, including codebases, infrastructure, and context (tickets, docs, conversations).
  • Design tasks from intermediate states of these environments, crafting prompts and defining what 'solved' means.
  • Ensure tasks are solvable by an AI agent.
  • Write tests that verify agent solutions, accepting all valid approaches and rejecting incorrect ones.
  • Iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining evaluations.

Benefits

  • Project-based work
  • Connect with leading tech companies
  • Flexible work schedule
  • Paid per accepted task
  • Potential effective hourly rate up to $150/hr
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service