We are seeking highly experienced software engineers (SR+) to evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role where you will be writing production code. Instead, you will be assessing a more nuanced aspect: whether the AI model 'thinks' like a great engineer. You will evaluate how AI coding agents behave in real-world scenarios, focusing on the sensibility of responses, the usefulness of preambles and reasoning, whether the output reflects strong engineering judgment, and the overall feel of the interaction for an experienced developer. This role emphasizes engineering 'taste' over mere syntax correctness. You will assess AI-generated coding interactions end-to-end, judging outputs for usefulness, high-level correctness, and alignment with strong engineering thought processes. You will also assess the quality of explanations and reasoning, distinguish between different levels of response quality, and provide clear, opinionated feedback on what worked, what didn't, and what felt 'off' or misleading. The goal is to help define what great looks like when interacting with tools like Cursor.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Part-time
Career Level
Senior
Education Level
No Education Listed