About The Position

We are sharing a specialised part-time consulting opportunity for experienced Senior Software Engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research. Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help researchers understand where advanced models succeed or fail throughout the software-development lifecycle. The work combines senior-level software engineering judgement with structured evaluation and benchmark design.

Requirements

  • Several years of professional software-engineering experience
  • At least 2 years of continuous full-time experience at a top-tier product company
  • Strong full-stack development capabilities
  • Experience building scalable, production-grade software
  • Deep understanding of software architecture and system design
  • Strong knowledge of debugging and code-quality assessment
  • Experience reviewing and improving complex production code
  • Proficiency in one or more of Python, JavaScript, C/C++, Java, Rust, or Go
  • Strong understanding of API design and production implementation
  • Familiarity with software monitoring and operational maintenance
  • Ability to reason across the complete software-engineering lifecycle
  • Strong analytical and problem-solving capabilities
  • Excellent written and verbal communication skills
  • Ability to produce clear and structured evaluation rationales
  • Comfortable collaborating with research and technical teams in a remote environment
  • Partial overlap with PST working hours is required

Nice To Haves

  • ReactJS experience is valuable for frontend-oriented assignments

Responsibilities

  • Curate high-quality code examples for model training and benchmarking
  • Develop precise reference solutions to software-engineering tasks
  • Correct and improve code across multiple programming languages
  • Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go
  • Maintain strong standards for correctness, clarity, and maintainability
  • Evaluate AI-generated code for technical correctness
  • Assess solutions for efficiency, scalability, and reliability
  • Identify implementation weaknesses and recurring error patterns
  • Review code quality against professional engineering standards
  • Provide structured rationales supporting evaluation decisions
  • Design mechanisms that automatically verify software-engineering solutions
  • Build agents capable of assessing code quality
  • Develop approaches for identifying common model failure patterns
  • Create deterministic or structured checks where appropriate
  • Support reliable evaluation across repeated coding tasks
  • Evaluate model capabilities across the software-development lifecycle
  • Assess reasoning around prototyping and architecture design
  • Review API design and production implementation decisions
  • Evaluate launch, experimentation, monitoring, and operational scenarios
  • Identify where models struggle with real-world engineering workflows
  • Collaborate with researchers and cross-functional technical teams
  • Contribute to datasets used for training and benchmarking
  • Help define engineering evaluation strategies and quality standards
  • Compare model performance against professional engineering expectations
  • Support iterative improvements to coding-focused evaluation systems

Benefits

  • No medical or paid-leave benefits are included under the contractor arrangement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service