About The Position

We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research. Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows.

Requirements

  • 3+ years of professional software-engineering experience
  • Strong full-stack development capabilities
  • Experience building scalable, production-grade software
  • Strong understanding of software architecture and system design
  • Deep knowledge of development, debugging, and code-quality assessment
  • Experience reviewing and improving complex software implementations
  • Proficiency in one or more of Python, JavaScript, Java, C++, Rust, or related languages
  • ReactJS, C, or Go experience may also be relevant to project assignments
  • Strong understanding of API design and production implementation
  • Familiarity with software monitoring and operational maintenance
  • Ability to reason across the complete software-engineering lifecycle
  • Strong analytical and problem-solving capabilities
  • Excellent written and verbal communication skills
  • Ability to provide clear, structured evaluation rationales
  • Comfortable collaborating remotely with research and technical teams
  • Candidates must be based in the United States, Canada, or eligible Western European (WEU) countries
  • Minimum commitment: 10 hours per week
  • Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

Nice To Haves

  • Source examples of WEU locations include Austria, Belgium, France, and Germany

Responsibilities

  • Curate high-quality code examples for model training and benchmarking
  • Develop precise solutions to software-engineering tasks
  • Correct and improve code across multiple programming languages
  • Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go
  • Maintain strong standards for correctness and maintainability
  • Evaluate AI-generated code for technical correctness
  • Assess solutions for efficiency, scalability, and reliability
  • Identify implementation weaknesses and recurring error patterns
  • Review code quality against professional engineering standards
  • Provide structured rationales supporting evaluation decisions
  • Build agents that assess code quality
  • Design mechanisms for automatically verifying software solutions
  • Identify recurring model-generated coding errors
  • Develop reliable checks for engineering tasks
  • Support reproducible evaluation across repeated assignments
  • Evaluate model capabilities across the software-development lifecycle
  • Assess reasoning around prototyping and architecture design
  • Review API design and production implementation decisions
  • Evaluate launch, experimentation, monitoring, and maintenance scenarios
  • Identify areas where models struggle with real-world engineering workflows
  • Collaborate with research and cross-functional technical teams
  • Contribute to datasets used for training and benchmarking
  • Help define engineering evaluation strategies
  • Compare model performance against professional engineering expectations
  • Support iterative improvements to coding-focused evaluation systems

Benefits

  • No medical or paid-leave benefits are included under the contractor arrangement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service