The Evaluation Tools group builds the internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behaviour. Model developers, researchers and QA engineers across Wayve depend on our tools to understand driving performance and scale evaluation to millions of scenarios, and every major model release runs through them. You'll join the Search & Agents squad in Sunnyvale. We build the search, scenario mining and agentic tooling that lets anyone find the right driving scenarios and turn them into tests without writing SQL. Our work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat. You'll work across the stack and own features end-to-end: talking to users, shaping ideas, building robust software, and validating impact. Evaluation is an evolving challenge, so you'll also have plenty of opportunity to define new projects as user needs emerge.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed