Full Stack Software Engineer, Evaluation Tools

WayveSunnyvale, CA
$209,000 - $266,000Hybrid

About The Position

The Evaluation Tools group builds internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behavior. Model developers, researchers, and QA engineers across Wayve depend on these tools to understand driving performance and scale evaluation to millions of scenarios. Every major model release runs through these tools. The role is within the Search & Agents squad in Sunnyvale, which builds search, scenario mining, and agentic tooling that allows users to find the right driving scenarios and turn them into tests without writing SQL. This work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat. The engineer will work across the stack, owning features end-to-end, including user interaction, idea shaping, building robust software, and validating impact. As evaluation is an evolving challenge, there will be opportunities to define new projects as user needs emerge.

Requirements

  • Strong development skills in Python, TypeScript, and JavaScript, with experience using React or similar front-end frameworks.
  • Strong SQL skills with a solid understanding of database design and optimisation.
  • Experience designing and building production-level fullstack systems and reliable, high-performance APIs.
  • Track record of building robust, maintainable systems and applying engineering best practices.
  • Excellent communication and collaboration skills, including working effectively across time zones.
  • Comfortable working independently in a fast-paced, high-context, ambiguous environment.
  • Power user of AI development tools and have a mindset for building AI-augmented workflows.

Nice To Haves

  • Experience with search, embeddings or retrieval systems.
  • Experience with the Databricks platform, and large-scale data processing with Spark or similar.
  • Experience with job orchestration frameworks such as Flyte.
  • Experience building agentic or LLM-powered developer tools (e.g. MCP servers).
  • Experience building tools for technical users (ML, robotics, infrastructure).
  • Experience with scientific data visualisation or complex data interfaces.

Responsibilities

  • Build fullstack features across search, scenario mining, agentic tooling, and the web applications our users work in to curate scenarios and build test suites.
  • Work directly with users to understand pain points, then define and measure what success looks like (adoption, time saved, reliability) before building.
  • Ship high-quality software with strong attention to reliability, performance, and maintainability.
  • Contribute to architectural decisions that scale across the team’s products.
  • Use AI tools to improve how we build, from speeding up development to refining team workflows and product quality.

Benefits

  • Competitive equity package
  • Hybrid working policy
  • Core working hours
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service