About The Position

We are sharing a specialised full-time consulting opportunity for experienced software engineers with strong Python development, debugging, version-control, technical documentation, and AI-assisted coding experience. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will design, implement, and review realistic multi-step software engineering tasks that test the capabilities of AI coding agents across Python development, environment setup, tooling, debugging, and technical problem-solving.

Requirements

  • At least 1 year of experience in software engineering, research engineering, or a related coding-intensive role
  • Strong hands-on Python scripting, implementation, and debugging skills
  • Experience developing clean, readable, and maintainable software
  • Everyday fluency with Git, IDEs, repositories, and standard software development workflows
  • Comfort configuring environments, dependencies, tooling, and validation processes
  • Strong technical writing and documentation skills
  • Ability to work independently through ambiguous and open-ended engineering problems
  • Reliable availability for approximately 35 hours per week

Nice To Haves

  • Experience using AI coding assistants, prompt engineering methods, or agent-based workflows
  • Previous work in AI training, model evaluation, or benchmark development
  • Background authoring technical tasks, reference solutions, or grading criteria
  • Familiarity with automated testing, CI/CD workflows, containers, or reproducible environments
  • Experience reviewing code or technical assignments created by other engineers
  • Knowledge of agentic AI systems and multi-step coding evaluations
  • Strong ability to identify edge cases, unintended shortcuts, and subtle implementation issues
  • Experience collaborating with AI research or evaluation teams

Responsibilities

  • Design realistic, multi-step software engineering challenges based on practical development workflows
  • Design technically demanding problems that require implementation, debugging, environment configuration, and analytical reasoning
  • Define clear requirements, constraints, expected outputs, and acceptance criteria
  • Ensure tasks assess genuine software engineering capability rather than superficial code generation
  • Build complete and verifiable reference solutions in Python
  • Create the supporting setup, dependencies, tests, and validation checks required for each task
  • Write clean, readable, and maintainable code
  • Confirm that solutions run reliably within the intended technical environment
  • Document implementation decisions and expected behaviour clearly
  • Use AI coding assistants and agent-based development tools within practical engineering workflows
  • Evaluate how frontier models approach complex coding and debugging tasks
  • Identify implementation errors, unsupported assumptions, inefficient approaches, and incomplete solutions
  • Analyse where AI agents succeed, struggle, or exploit unintended shortcuts
  • Document failure patterns and provide evidence supporting evaluation conclusions
  • Review tasks and reference solutions created by other software engineering specialists
  • Assess clarity, correctness, difficulty, reproducibility, and technical fairness
  • Identify ambiguous instructions, hidden assumptions, grading gaps, and environment issues
  • Provide actionable feedback that improves task quality and benchmark reliability
  • Collaborate closely with researchers and fellow task authors

Benefits

  • Competitive hourly compensation
  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service