About The Position

Siri's quality signal drives every model and product decision before a release ships, and that signal is only as good as the tests and scenarios behind it. The Test Development Tooling team, part of the Siri's Agentic Evaluation Engineering organization, is reimagining how those scenarios are created and maintained. We build tools and AI-powered automation that free our engineers to focus on the quality and strategy of evaluation rather than the manual work of authoring tests. We're looking for an inspiring, hands-on engineering leader to run this team and set its technical direction. The ideal candidate is technically strong player-coach that will provide technical direction for the team.

Requirements

  • 5+ years of experience developing production software (e.g., Swift, Python, Java, or Go).
  • 3+ years of engineering management experience.
  • Strong technical leadership and the ability to set architectural direction and inspire a team.
  • Experience with test infrastructure, CI/CD, or evaluation/ML platforms.
  • Proactive and self-motivated, with demonstrated creative and critical-thinking abilities.
  • Experience building or leading teams that apply LLMs / AI agents to real engineering workflows (e.g., automated test generation, code generation, or agentic pipelines with human-in-the-loop review).
  • Depth in one or more of: distributed systems and backend services (REST/gRPC), cloud infrastructure (Docker/Kubernetes), data platforms, or large-scale test and release infrastructure.
  • Demonstrated cross-functional leadership and the ability to drive coverage and quality decisions with senior stakeholders.
  • Experience hiring and growing a team, including senior engineers.
  • Excellent spoken and written communication skills.
  • Comfort operating in ambiguity and reshaping team scope as the domain matures.

Nice To Haves

  • Familiarity with Apple's testing frameworks (XCTest and related) and with ML/LLM evaluation concepts.

Responsibilities

  • Lead and grow a team of engineers building AI agent-based scenario authoring and selection tools.
  • Set the team's technical direction and roadmap.
  • Own the systems that generate, validate, ground, and manage evaluation scenarios through their full lifecycle.
  • Drive coverage traceability and gap analysis so that every product capability maps to a scenario, and every gap is visible.
  • Deliver and evolve coverage, reliability, and health dashboards and release-readiness signals that leadership relies on for ship decisions.
  • Collaborate across the evaluation organization and Apple to integrate and ground scenarios end to end at scale.
  • Partner with QE engineers, feature teams, architecture leads, and design stakeholders to set and hold a defensible quality bar.
  • Recruit, mentor, and develop engineers, including senior ICs, with a focus on engineering rigor, technical depth, and clear ownership.
  • Operate effectively in ambiguity, reshaping team scope as the evaluation platform and the underlying Siri architecture evolve.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service