About The Position

We are sharing a specialised full-time consulting opportunity for researchers with strong computational, experimental-design, data-analysis, and scientific reasoning experience across STEM, computational social science, or computational humanities disciplines. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected researchers will translate real research methods—including study design, hypothesis testing, simulation, modelling, and rigorous evaluation—into complex multi-step tasks requiring coding, data analysis, and carefully supported conclusions.

Requirements

  • At least 1 year of experience in an active academic, industry, government, or national-laboratory research role
  • An MSc or PhD in a STEM field, computational social science, computational humanities, or another research-intensive discipline
  • Significant experience using Python for analysis, simulation, modelling, or data pipelines
  • Strong grounding in experimental design, hypothesis testing, and rigorous evaluation
  • Experience interpreting complex datasets and producing evidence-based conclusions
  • Working familiarity with Git, IDEs, and Jupyter or Colab notebooks
  • Strong written communication and technical documentation skills
  • Ability to work independently through ambiguous and open-ended research problems
  • Reliable availability for approximately 35 hours per week

Nice To Haves

  • Experience in AI training, model evaluation, or benchmark development
  • Background authoring technical tasks, reference solutions, or grading rubrics
  • Familiarity with machine learning, statistical modelling, or scientific computing
  • Experience designing reproducible experiments and validating analytical pipelines
  • Knowledge of research-quality assurance, peer review, or methodological auditing
  • Familiarity with agentic AI systems and multi-step model evaluations
  • Experience reviewing code, notebooks, or technical analyses prepared by other researchers
  • Strong ability to identify subtle methodological flaws and unsupported conclusions

Responsibilities

  • Transform real research workflows into challenging, multi-step benchmark tasks
  • Develop assignments involving study design, hypothesis testing, simulation, modelling, or data analysis
  • Create tasks that assess genuine scientific reasoning rather than surface-level pattern matching
  • Ensure each task includes realistic assumptions, constraints, datasets, and evaluation objectives
  • Implement complete reference solutions using Python and notebook environments
  • Build reproducible analyses, simulations, models, or data-processing pipelines
  • Validate calculations, code, intermediate outputs, and final conclusions
  • Document methodologies clearly enough for independent review and reproduction
  • Define what distinguishes rigorous research reasoning from plausible but unsupported analysis
  • Develop reference answers, grading criteria, and structured evaluation guidelines
  • Identify required methodological steps, valid alternative approaches, and material errors
  • Ensure grading standards reflect the quality expected from an experienced researcher
  • Evaluate AI-generated attempts across computational research tasks
  • Identify methodological errors, unsupported claims, statistical weaknesses, and coding issues
  • Assess whether conclusions follow logically from the evidence and analysis
  • Provide clear written feedback explaining errors and recommended improvements
  • Work closely with researchers, task authors, and fellow subject-matter experts
  • Compare evaluation decisions and support consistent benchmark standards
  • Incorporate feedback into task design, reference solutions, and scoring criteria
  • Surface recurring model failure patterns and opportunities for stronger evaluation coverage

Benefits

  • Competitive hourly compensation
  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service