We're looking for a software engineer to help build and evaluate CUA (Computer Use Agents) — models trained to operate real software and complete computer-based tasks the way a skilled engineer would. You'll own complex, long-running technical workflows end to end — implementation, validation, documentation, and follow-through — and design and maintain the tasks, benchmarks, and test suites that measure CUA capability as the underlying models, APIs, and infrastructure keep evolving. You hold an exacting bar for code quality: you love writing tests, you're fluent with deployment strategies like canary releases and rollbacks, and you bring a strong sense of risk management. You have a sharp eye for weak implementations and give (and take) blunt, direct feedback. You're also fluent with AI coding assistants like Claude Code, Codex, or Cursor, with your own well-developed best practices for using them, and you're proficient in Rust, Python, and/or TypeScript. Above all, you bring strong debugging and experimental discipline — the ability to investigate discrepancies across code, configuration, infrastructure, and results within the CUA stack, and turn ambiguous findings into reproducible conclusions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed