About The Position

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria within realistic simulated environments. This role involves building realistic developer environments (a virtual company with codebase, infrastructure, and context), designing tasks from intermediate states of these environments, writing tests to verify agent solutions, and iterating on tasks and tests based on QA feedback. This is not data labeling, prompt engineering, or writing code from scratch.

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+
  • A Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems or other related fields.
  • Bachelor’s degree is accepted if only candidate has 5 years of experience in the field.
  • Minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles

Nice To Haves

  • Deep understanding of where AI models fail and what scenarios reveal the difference between a good and a bad solution.

Responsibilities

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Benefits

  • Paid per accepted task.
  • Rate depends on qualification tier and efficiency - up to the equivalent of $30/hr.
  • Faster pace raises effective hourly rate.
  • Part-time, remote, freelance project.
  • Fits around primary professional or academic commitments.
  • Work on advanced AI projects and gain valuable experience.
  • Enhances portfolio.
  • Influence how future AI models understand and communicate in your field of expertise.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service