About The Position

We are sharing a specialised full-time consulting opportunity for experienced data scientists and quantitative analysts with strong expertise in statistical analysis, data cleaning, method comparison, reproducible research, and evidence-based reporting. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will create realistic data-analysis challenges, develop reproducible reference notebooks, evaluate model-generated analyses, and identify where statistical reasoning, interpretation, or reporting falls short of professional standards.

Requirements

  • At least 1 year of experience in data science, quantitative analysis, research engineering, or another research-intensive analytical role
  • Deep hands-on experience with data cleaning, statistical correlation, hypothesis testing, and interpretation
  • Strong proficiency in Python, including pandas, NumPy, or comparable analytical libraries
  • Experience using Jupyter Notebook or Google Colab for analysis and reporting
  • Working familiarity with Git and reproducible analytical workflows
  • Ability to communicate complex quantitative findings clearly to technical and non-technical decision-makers
  • Strong attention to detail and confidence working through ambiguous, open-ended problems
  • Reliable availability for approximately 35 hours per week

Nice To Haves

  • Experience in AI training, model evaluation, or benchmark development
  • Background authoring analytical tasks, reference solutions, or grading rubrics
  • Familiarity with anomaly detection, experimental design, or comparative model evaluation
  • Experience conducting manual spot checks and validating automated analyses
  • Knowledge of statistical modelling, machine learning, or scientific computing
  • Familiarity with agentic AI systems and multi-step model evaluations
  • Experience reviewing notebooks, code, or analyses prepared by other professionals
  • Strong ability to identify subtle statistical errors and unsupported conclusions

Responsibilities

  • Create realistic analytical tasks based on professional data science and quantitative research workflows
  • Develop assignments involving messy data, anomaly detection, correlation analysis, hypothesis testing, and method comparison
  • Design complex, multi-step problems requiring statistical judgment and careful interpretation
  • Ensure tasks include realistic constraints, datasets, assumptions, and decision-making objectives
  • Complete reference analyses using Jupyter Notebook or Google Colab
  • Build clear and reproducible workflows using Python, pandas, NumPy, and related libraries
  • Document data-cleaning decisions, calculations, statistical methods, and analytical conclusions
  • Validate intermediate results, spot checks, visualisations, and final recommendations
  • Design fair comparisons between analytical models, algorithms, or statistical approaches
  • Evaluate performance using appropriate metrics, manual checks, and sensitivity analyses
  • Identify methodological trade-offs, limitations, and sources of uncertainty
  • Produce recommendations supported by transparent quantitative evidence
  • Review model-generated analyses for statistical accuracy, methodological rigour, and sound interpretation
  • Verify whether calculations, correlations, hypotheses, and conclusions are supported by the data
  • Identify coding errors, unsupported assumptions, misleading summaries, and analytical shortcuts
  • Explain where and why model outputs fail to meet professional data-analysis standards
  • Work closely with researchers, task authors, and fellow quantitative specialists
  • Compare evaluation decisions to maintain consistent benchmark standards
  • Refine tasks, reference notebooks, and grading criteria based on testing outcomes
  • Document recurring model weaknesses and opportunities for stronger evaluation coverage

Benefits

  • Competitive hourly compensation
  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service