Remote | QA Test Engineer — $55–$85/hour

24-MagNew York, NY
Remote

About The Position

We are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.

Requirements

  • At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical role
  • Demonstrated experience designing test cases and quality-review processes
  • Strong end-to-end debugging skills across complex technical systems
  • Working proficiency in Python and Git
  • Comfort navigating unfamiliar codebases, repositories, and execution environments
  • Exceptional attention to detail and strong written documentation habits
  • Ability to identify ambiguity, edge cases, hidden assumptions, and quality gaps
  • Capacity to work independently through open-ended technical problems
  • Reliable availability for approximately 35 hours per week

Nice To Haves

  • Experience with AI training, model evaluation, or quality review of AI-generated work
  • Familiarity with agentic systems and multi-step AI benchmarks
  • Background testing machine learning, research, or data-processing workflows
  • Experience developing automated test suites or validation scripts
  • Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
  • Experience reviewing reference solutions, grading logic, or technical rubrics
  • Knowledge of adversarial testing, failure-mode analysis, or benchmark design
  • Prior collaboration with AI research or evaluation teams

Responsibilities

  • Create comprehensive test cases confirming that benchmark tasks function as intended
  • Design positive, negative, boundary, and edge-case tests
  • Validate task requirements, expected outputs, reference solutions, and grading logic
  • Identify scenarios that may produce incorrect or misleading evaluation results
  • Ensure tests measure the intended technical capability accurately
  • Review complex multi-step tasks and reference solutions before finalisation
  • Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteria
  • Run tasks independently to confirm reproducibility and expected behaviour
  • Assess whether grading standards are clear, fair, and technically defensible
  • Provide actionable feedback to task authors and researchers
  • Investigate failures across Python scripts, test harnesses, repositories, and task environments
  • Diagnose unexpected behaviour within unfamiliar codebases
  • Reproduce reported issues and isolate their underlying causes
  • Correct or document environment, dependency, logic, and validation problems
  • Use Git-based workflows to support structured review and collaboration
  • Develop practical checklists and repeatable review procedures for benchmark quality
  • Improve consistency across task validation, testing, and approval workflows
  • Document findings clearly so authors can resolve issues efficiently
  • Track recurring defects and recommend preventive quality measures
  • Collaborate closely with researchers, task authors, and other technical reviewers
  • Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
  • Identify cases where models can receive credit without completing the intended reasoning or technical work
  • Test whether benchmark tasks remain robust across alternative approaches
  • Strengthen evaluation criteria to maintain reliable and meaningful benchmark scores
  • Distinguish valid solution diversity from unintended task exploitation

Benefits

  • Competitive hourly compensation
  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service