As part of the Siri organization, you will build the systems and tooling that make evaluation a first-class part of how Siri is developed — not an after-the-fact check — spanning human evaluation, real user feedback, reward and alignment signals, and data science rigor across iOS, iPadOS, macOS, watchOS, and visionOS. This is a rare opportunity to work at the intersection of software engineering and rigorous evaluation science — applying machine learning engineering techniques, from model evaluation to reward modeling, to build the infrastructure that keeps this quality signal trustworthy. What you build will directly shape the direction of one of the world's most widely used assistants. In this role you'll contribute across several interconnected work streams spanning evaluation quality, reward/alignment signals, and data science. A core part of the job is bringing "evals first" thinking to the team — building tooling and harnesses grounded in real workflows, and designing for observability and reproducibility so quality can be measured clearly and issues caught early. This is a largely unexplored space with few established playbooks, so being self-driven is a must — you'll define your own path as much as execute one. Scope and priorities will evolve, and we're looking for someone who moves fluidly across these areas, bringing strong software engineering fundamentals with enough ML/LLM depth to build and ship AI-facing tooling.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior