About The Position

We are sharing a specialised part-time consulting opportunity for experienced Senior Software Engineers with tech-lead-level expertise to contribute to advanced LLM evaluation, repository validation, and realistic software-engineering benchmark development. Selected professionals will work with high-quality public repositories to develop and validate verifiable software-engineering tasks, configure reproducible development environments, analyse open-source issues, assess test quality, and evaluate how advanced language models perform on realistic bug-fixing and code-modification scenarios.

Requirements

  • Senior or tech-lead-level software-engineering experience
  • Strong expertise in at least one of Python, JavaScript, Java, Go, Rust, C, C++, C#, or Ruby
  • Comfortable working across unfamiliar and complex codebases
  • Strong proficiency with Git and repository-based development workflows
  • Practical Docker and environment-setup experience
  • Ability to run, modify, debug, and test production-quality software locally
  • Strong understanding of software testing and unit-test design
  • Ability to evaluate code quality, test coverage, and implementation correctness
  • Experience working with high-quality public repositories
  • Familiarity with widely used repositories with 500+ stars
  • Strong debugging and issue-triage capabilities
  • Ability to identify technically challenging software-engineering tasks
  • Clear written communication and technical reasoning skills
  • Comfortable collaborating remotely with research and engineering teams

Nice To Haves

  • Experience contributing to or evaluating open-source software is advantageous
  • Previous LLM research or evaluation experience is advantageous but not required
  • Experience with developer tools or software-automation agents is also beneficial

Responsibilities

  • Repository Analysis & Issue Triage: Analyse issues across well-maintained public software repositories, Identify technically meaningful problems suitable for LLM evaluation, Triage issues by complexity, reproducibility, and engineering relevance, Navigate repository history to understand implementation context, Prioritise problems that require substantive software-engineering reasoning
  • Repository Setup & Environment Validation: Configure complex repositories for reliable local execution, Build reproducible development and testing environments, Containerise projects using Docker where appropriate, Resolve dependencies and environment-specific configuration issues, Validate that repositories can be consistently executed and tested
  • Code Modification & LLM Evaluation: Modify and run real-world codebases locally, Evaluate model performance on bug-fixing and implementation tasks, Assess whether generated solutions correctly address repository issues, Identify failure patterns across different programming languages and task types, Compare generated solutions against professional engineering expectations
  • Testing & Verification Quality: Evaluate unit-test coverage and overall test quality, Determine whether existing tests adequately validate intended behaviour, Identify missing edge cases and weak verification mechanisms, Develop or refine tests where stronger validation is required, Ensure evaluation tasks have objective and reproducible outcomes
  • Research Collaboration & Technical Leadership: Collaborate with research teams on repository and task selection, Identify software-engineering problems that remain challenging for LLMs, Contribute insight into benchmark difficulty and dataset coverage, Support expansion across programming languages and task complexity, Provide technical leadership or guidance to junior engineers where required

Benefits

  • No medical or paid-leave benefits are included under the contractor arrangement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service