About The Position

We are sharing a specialised part-time consulting opportunity for experienced Machine Learning professionals with strong hands-on expertise in experiment design, model selection, hyperparameter tuning, data quality, and rigorous evaluation methodology. This role focuses on reviewing applied machine learning tasks for technical correctness, methodological rigour, and reproducibility. Selected experts will assess experimental design, modelling decisions, validation practices, metrics, and supporting evidence while identifying issues such as data leakage, metric gaming, or unreliable conclusions.

Requirements

  • 3+ years of hands-on applied or experimental machine learning experience
  • Strong practical experience with experiment design, model selection, hyperparameter tuning, and evaluation methodology
  • Deep understanding of data leakage, metric gaming, train/test methodology, and cross-validation hygiene
  • Proficiency with standard ML frameworks such as PyTorch, TensorFlow, scikit-learn, or XGBoost
  • Strong ability to critique machine learning claims against experimental evidence
  • Comfortable reproducing results and diagnosing discrepancies
  • Strong quantitative and analytical judgement
  • Clear written communication and ability to provide precise technical feedback

Nice To Haves

  • Kaggle, ML competition, or benchmark experience is advantageous
  • Graduate research or publication experience in applied machine learning is preferred
  • Previous task-grading, technical peer-review, or ML evaluation experience is advantageous

Responsibilities

  • Evaluate applied machine learning experiments for methodological soundness
  • Review experimental hypotheses, assumptions, and modelling choices
  • Assess whether experiments are appropriately designed to answer the stated question
  • Identify weaknesses that could invalidate or distort conclusions
  • Apply practical judgement grounded in hands-on experimental ML experience
  • Review model-selection decisions and supporting rationale
  • Evaluate hyperparameter-tuning approaches and search strategies
  • Assess whether model comparisons are fair and methodologically appropriate
  • Identify overfitting, cherry-picking, or poorly justified modelling choices
  • Determine whether conclusions are supported by experimental results
  • Evaluate train/test splits, cross-validation strategies, and validation procedures
  • Identify inappropriate data partitioning or evaluation practices
  • Review whether datasets and experimental protocols support reliable generalisation
  • Detect contamination between training, validation, and test data
  • Assess whether evaluation procedures match the underlying ML problem
  • Review datasets and preprocessing workflows for potential quality issues
  • Identify data leakage, target leakage, or unintended information exposure
  • Assess feature-engineering and preprocessing decisions
  • Detect methodological shortcuts that could inflate reported performance
  • Evaluate whether data handling supports reliable experimentation
  • Review metric selection against the task objective
  • Assess whether reported metrics appropriately capture model performance
  • Identify metric gaming, misleading optimisation targets, or incomplete evaluation
  • Review performance comparisons and supporting statistical evidence
  • Determine whether claimed improvements are meaningful and defensible
  • Reproduce or validate experimental results where required
  • Assess whether reported findings can be independently reproduced
  • Review code, configuration, experimental settings, and supporting evidence
  • Identify inconsistencies between claims and observed results
  • Determine whether conclusions follow logically from the available evidence
  • Work with experiments built using standard ML frameworks and libraries
  • Review implementations involving PyTorch, TensorFlow, scikit-learn, and XGBoost
  • Evaluate model-training and experimentation workflows
  • Identify implementation choices that may compromise experimental validity
  • Apply framework-specific knowledge when assessing technical quality
  • Review applied ML tasks and benchmark-style challenges
  • Assess whether tasks measure the intended modelling capability
  • Evaluate competition-style experimental approaches where relevant
  • Identify task-design issues that could reward shortcuts rather than genuine modelling quality
  • Apply experience from benchmarking or competitive ML environments where applicable
  • Assess assigned ML tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Identify specific methodological or implementation evidence behind each judgement
  • Apply grading standards consistently across assignments
  • Distinguish genuine methodological problems from reasonable alternative approaches

Benefits

  • Flexible scheduling based on project requirements
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service