About The Position

We are seeking highly experienced software engineers (Senior level and above) to evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role where you will be writing production code. Instead, the focus is on assessing whether the AI model 'thinks' like a great engineer. You will evaluate how AI coding agents behave in real-world scenarios, focusing on the usefulness of responses, the quality of preambles and reasoning, whether the output reflects strong engineering judgment, and the overall feel of the interaction for an experienced developer. This role emphasizes engineering 'taste' over mere syntax correctness. You will assess AI-generated coding interactions end-to-end, judging outputs for usefulness, high-level correctness, and alignment with strong engineering thought processes. You will also assess the quality of explanations and reasoning, distinguish between different levels of response quality, and provide clear, opinionated feedback on what worked, what didn't, and what felt 'off' or misleading. The goal is to help define what 'great' looks like when interacting with tools like Cursor. The role requires engineers who can make subjective but rigorous judgments about whether an AI's response feels like something a strong engineer would say, if an explanation is helpful or just technically correct, if the model is guiding the user well, and whether the interaction builds or erodes trust.

Requirements

  • Staff / Principal-level engineer (or equivalent experience)
  • Strong background in TypeScript / JavaScript or Python
  • Hands-on experience using OpenAI Codex
  • Hands-on experience using Claude Code
  • Hands-on experience using Cursor
  • Deep familiarity with modern AI-assisted dev workflows
  • Able to evaluate code without needing to fully execute or deeply review every line
  • Comfortable giving direct, opinionated feedback
  • High bar for what “good engineering” looks like

Nice To Haves

  • Experience with tools like Cursor or similar AI-first IDEs
  • Prior exposure to prompt design or evaluation workflows
  • Experience mentoring senior engineers or defining engineering standards

Responsibilities

  • Evaluate AI-generated coding interactions end-to-end
  • Judge whether outputs are useful, correct (at a high level), and aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning, not just code
  • Distinguish between different levels of response quality (e.g. what makes something a 2 vs 4)
  • Provide clear, opinionated feedback on what worked, what didn’t, and what felt “off” or misleading
  • Help define what great looks like when interacting with tools like Cursor
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service