About The Position

We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation. This role focuses on reviewing software-engineering benchmark tasks for correctness, reproducibility, and grading integrity. Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design.

Requirements

  • 3+ years of professional software engineering experience
  • Demonstrated open-source contribution or maintainer experience, such as merged pull requests, committer responsibilities, or maintainer roles
  • Strong ability to review repository-level software changes
  • Experience auditing reference patches, test runners, and automated test suites
  • Comfortable evaluating Docker-based isolation and reproducible development environments
  • Strong understanding of software testing, debugging, and code-review practices
  • Ability to identify answer leakage, reward hacking, or other benchmark-integrity issues
  • Strong proficiency in Python
  • Professional fluency in at least one additional language such as Java, Go, TypeScript, or C++

Nice To Haves

  • Familiarity with SWE-Bench Verified or similar repository-level software engineering benchmarks is preferred
  • Maintainer or contributor history with established Python open-source projects is highly valued
  • Previous code-review, software evaluation, or task-grading experience is advantageous
  • Strong written communication and ability to provide precise technical feedback

Responsibilities

  • Repository-Level Code Review: Review software engineering tasks built around real code repositories, assess whether task requirements are technically clear, complete, and reproducible, evaluate repository state, dependencies, configuration, and expected behaviour, identify ambiguities or implementation issues that could affect task validity, and apply practical engineering judgement to realistic codebase-level problems.
  • Reference Patch Auditing: Review reference patches for correctness and completeness, determine whether proposed solutions appropriately address the underlying software issue, identify unintended behavioural changes, incomplete fixes, or unsupported assumptions, compare reference implementations against task requirements and expected outcomes, and assess whether alternative valid implementations are treated fairly.
  • Test Harness & Grading Review: Audit test runners and automated evaluation logic, assess whether tests accurately measure the intended behaviour, identify missing coverage, brittle assertions, or grading inconsistencies, verify that evaluation criteria appropriately distinguish correct from incorrect solutions, and review benchmark tasks for reliable and repeatable scoring.
  • Reproducibility & Environment Validation: Evaluate whether tasks can be reproduced consistently across clean environments, review dependency installation, build processes, configuration, and runtime requirements, assess Docker-based isolation and containerised execution, identify environmental dependencies or hidden assumptions affecting reproducibility, and verify that tasks execute reliably under their intended setup.
  • Benchmark Integrity: Identify potential answer leakage, unintended shortcuts, or reward-hacking opportunities, evaluate whether benchmark structure exposes information that makes tasks artificially easy, review task and grading design for loopholes or exploitable behaviours, assess whether successful completion genuinely demonstrates the intended engineering capability, and recommend improvements where benchmark integrity is compromised.
  • Software Testing & Debugging: Investigate failing or inconsistent benchmark tasks, review stack traces, logs, test failures, and repository behaviour, identify root causes of technical issues, distinguish task defects from legitimate implementation failures, and assess whether debugging and validation processes follow sound engineering practices.
  • Open-Source Engineering: Apply experience from contributing to or maintaining open-source software, evaluate repository conventions, contribution patterns, and realistic development workflows, review patches with the perspective of an experienced contributor or maintainer, assess whether proposed changes would meet reasonable code-review expectations, and apply practical judgement derived from real-world pull request and repository experience.
  • Multi-Language Code Evaluation: Review software written in Python, evaluate tasks involving at least one additional ecosystem such as Java, Go, TypeScript, or C++, assess code structure, tests, implementation choices, and repository conventions across languages, identify language-specific implementation or testing issues, and apply consistent engineering standards across different technology stacks.
  • Rubric-Based Evaluation: Assess benchmark tasks against structured technical criteria, provide clear written explanations supporting evaluation decisions, reference specific code, tests, patches, or execution behaviour when identifying issues, apply grading standards consistently across assignments, and distinguish substantive benchmark defects from minor implementation differences.

Benefits

  • Flexible scheduling based on project requirements
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service