Full Stack Software Engineer, Evaluation Tools

Wayve•Sunnyvale, CA
•$209,700 - $311,400•Hybrid

About The Position

The Evaluation Tools group builds internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behavior. Model developers, researchers, and QA engineers across Wayve depend on these tools to understand driving performance and scale evaluation to millions of scenarios. Every major model release runs through these tools. The role is within the Search & Agents squad, which builds search, scenario mining, and agentic tooling that allows users to find the right driving scenarios and turn them into tests without writing SQL. This work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat. The engineer will work across the stack, owning features end-to-end, including user interaction, idea shaping, software building, and impact validation. As evaluation is an evolving challenge, there will be opportunities to define new projects as user needs emerge.

Requirements

  • Strong development skills in Python, TypeScript, and JavaScript, with experience using React or similar front-end frameworks.
  • Strong SQL skills with a solid understanding of database design and optimisation.
  • Experience designing and building production-level fullstack systems and reliable, high-performance APIs.
  • Track record of building robust, maintainable systems and applying engineering best practices.
  • Excellent communication and collaboration skills, including working effectively across time zones.
  • Comfortable working independently in a fast-paced, high-context, ambiguous environment.
  • A power user of AI development tools and a mindset for building AI-augmented workflows.

Nice To Haves

  • Experience with search, embeddings or retrieval systems.
  • Experience with the Databricks platform, and large-scale data processing with Spark or similar.
  • Experience with job orchestration frameworks such as Flyte.
  • Experience building agentic or LLM-powered developer tools (e.g. MCP servers).
  • Experience building tools for technical users (ML, robotics, infrastructure).
  • Experience with scientific data visualisation or complex data interfaces.

Responsibilities

  • Build fullstack features across search, scenario mining, and agentic tooling.
  • Talk to model developers and researchers, and define how success will be measured before building.
  • Take the scenario mining framework from prototype to production scale.
  • Extend agent tooling so agents can run evaluation workflows end to end.
  • Work with large-scale data systems (Databricks, Spark) in an established codebase.
  • Own the reliability and maintainability of shipped software.
  • Build fullstack features across search, scenario mining, agentic tooling, and web applications for curating scenarios and building test suites.
  • Work directly with users to understand pain points, then define and measure success (adoption, time saved, reliability) before building.
  • Ship high-quality software with strong attention to reliability, performance, and maintainability.
  • Contribute to architectural decisions that scale across the team’s products.
  • Use AI tools to improve how we build, from speeding up development to refining team workflows and product quality.

Benefits

  • Salaries benchmarked against the market annually
  • Meaningful equity, sharing in the ownership and long term success of Wayve
  • Relocation support and visa sponsorship where applicable
  • Hybrid working, core hours and the chance to work hands on in vehicle workshops and labs
  • Learning and development budgets with support for training, conferences and growth
  • Comprehensive benefits including health insurance, dental, enhanced maternity and paternity leave, retirement or pension where applicable, access to therapists, wellbeing partnerships, team socials and more
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service