The Evaluation Tools group builds internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behavior. Model developers, researchers, and QA engineers across Wayve depend on these tools to understand driving performance and scale evaluation to millions of scenarios. Every major model release runs through these tools. The Search & Agents squad builds the search, scenario mining, and agentic tooling that allows users to find the right driving scenarios and turn them into tests without writing SQL. This work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat. The role involves working across the stack and owning features end-to-end: talking to users, shaping ideas, building robust software, and validating impact. There will be opportunities to define new projects as user needs emerge.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed