About The Position

We are sharing a specialised part-time consulting opportunity for experienced machine learning engineers with hands-on experience using AI coding agents and building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products. This sprint-based role supports an advanced AI research initiative focused on evaluating frontier coding models through realistic machine learning engineering workflows. Selected professionals will use AI coding agents to complete technical tasks, review model-generated implementations, identify bugs and failure modes, and compare how different models perform across practical ML engineering scenarios.

Requirements

  • At least 2 years of professional machine learning engineering experience
  • Experience building production ML systems, AI-powered applications, or model-serving infrastructure
  • Hands-on experience with model training, inference, deployment, or MLOps
  • Experience developing LLM applications or integrating foundation models into production systems
  • Regular use of AI coding agents within software or machine learning development workflows
  • Strong ability to evaluate model-generated code and technical implementation decisions
  • Excellent debugging, analytical reasoning, and written communication skills
  • Ability to work efficiently within short, intensive project sprints
  • A degree in computer science, machine learning, artificial intelligence, software engineering, or a related technical discipline may be helpful
  • Equivalent professional experience building and deploying production ML systems may also be considered
  • Practical engineering depth is particularly important for this engagement

Nice To Haves

  • Experience with Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or comparable AI coding tools
  • Production experience deploying machine learning models
  • Familiarity with model-serving architectures and inference optimisation
  • Experience building LLM-powered applications or agentic systems
  • Knowledge of MLOps, deployment pipelines, monitoring, or model infrastructure
  • Experience evaluating generated code across multiple AI coding systems
  • Previous exposure to AI evaluation, benchmark development, or structured technical review
  • Advanced study in machine learning or computer science may strengthen an application

Responsibilities

  • Review complex machine learning and AI engineering tasks completed with frontier coding agents
  • Evaluate implementations involving model training, inference systems, MLOps, and LLM applications
  • Assess technical correctness, architecture choices, implementation quality, and engineering trade-offs
  • Apply professional ML engineering judgment to realistic production-oriented scenarios
  • Use AI coding agents as part of hands-on technical workflows
  • Evaluate how effectively coding models interpret requirements and implement solutions
  • Identify bugs, incomplete implementations, edge cases, and unexpected behaviour
  • Assess where models require additional prompting, correction, or manual engineering intervention
  • Identify performance issues, reliability problems, and model failure modes
  • Review generated code for maintainability, correctness, and practical usability
  • Evaluate whether implementations would function appropriately in realistic ML environments
  • Document technical strengths, weaknesses, and important implementation risks
  • Compare outputs produced by multiple frontier coding models
  • Assess differences in implementation strategy, code quality, technical reasoning, and reliability
  • Determine which approaches best satisfy task requirements
  • Provide clear written assessments explaining relevant engineering trade-offs

Benefits

  • Sprint-based project with task-based compensation
  • Compensation is $400 per accepted task
  • Typical tasks require approximately 2–3 hours after ramp-up
  • Weekly payments via Stripe or Wise
  • Projects may be extended, shortened, or adjusted depending on scope and performance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service