About The Position

We are hiring senior full-stack engineers to help build and refine next-generation coding environments that evaluate AI agents. This work directly supports the development of robust grading systems that accurately measure how effectively an AI agent solves complex coding tasks. The goal is to create secure evaluation suites that cannot be bypassed or easily cheated by the models. Throughout this long-term engagement, you will work remotely with real website clones to identify bugs and write technical specifications. You will build comprehensive automated test suites designed to grade AI-generated code. This involves analyzing how an agent interacts with a given task and ensuring the grading logic is watertight. The process requires deep technical scrutiny and creative problem-solving to account for unpredictable AI behavior.

Requirements

  • Senior-level experience in full-stack software engineering
  • Deep professional expertise building web applications with React
  • Strong background in automated testing and technical specification writing
  • Available for an ongoing, high-commitment technical engagement

Nice To Haves

  • Strong background in web security
  • Strong background in QA automation
  • Strong background in AI evaluation

Responsibilities

  • Interact with real website clones to identify bugs and workflow gaps
  • Write detailed technical specifications for complex AI coding tasks
  • Build robust automated test suites to securely grade AI agent performance
  • Ensure grading logic is watertight and resistant to AI bypassing or shortcuts
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service