We’re looking for highly experienced software engineers (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role where you will be writing production code. Instead, you’ll be evaluating something harder: whether the model thinks like a great engineer. You will assess how AI coding agents behave in real-world scenarios, focusing on whether the response makes sense, whether the preamble and reasoning are useful, whether the output reflects strong engineering judgment, and whether the interaction feels right to an experienced developer. This role is about engineering taste — not syntax correctness.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Part-time
Career Level
Senior
Education Level
No Education Listed
Number of Employees
11-50 employees