About The Position

The Apple Intelligence Platform Experience Validation team builds the tooling and automation that keeps Apple Intelligence features high-quality before they ship. We are looking for a Senior SDET to lead the design and implementation of automated model evaluation: standing up LLM-as-a-judge in existing and new pipelines, and building the infrastructure that catches model regressions before they reach human evaluation or the live on population. This is a hands-on, senior individual-contributor role. You will own eval automation as a discipline across the team, partnering with modeling, framework, and infrastructure teams to make model quality a first-class, continuously measured signal.

Requirements

  • BS in Computer Science, Mathematics, or a related field (or equivalent practical experience)
  • Three years of relevant industry experience in test automation, software development, or related areas.

Nice To Haves

  • Strong practical knowledge of Python, including data-pipeline fluency (JSON/YAML, REST APIs).
  • Hands-on experience with LLM-as-a-judge evaluation and rubric design, or a strong demonstrated ability to ramp into it quickly.
  • Strong software engineering fundamentals — able to define atomic, composable components and build maintainable pipelines and tooling, not just scripts.
  • Strong debugging and triage skills; able to separate genuine regressions from infrastructure or rubric noise.
  • Strong knowledge of the software development lifecycle, testing methodologies, and QA processes.
  • Excellent written and verbal communication; able to document clearly and describe quality signal to modeling and leadership audiences.
  • Ability to lead work across varying priorities and partner multi-functionally with modeling, framework, and infrastructure teams.
  • Experience building on-device tooling and device/model eval infrastructure.
  • Experience integrating with CI/CD and job orchestration systems, and comfort deploying tooling as reusable libraries.
  • Familiarity with generative model behavior — image generation, NLP, or LLM output evaluation.
  • Experience curating and reasoning about large datasets; comfort manually inspecting data (Jupyter or similar) to build intuition and drive next steps.
  • Awareness of dataset bias and fairness considerations in evaluation.
  • Experience with database/query tooling (e.g., SQL) and dashboards/visualization for reporting quality trends.
  • Experience with Xcode is a bonus.

Responsibilities

  • Build and maintain model level, component or end-to-end evaluation coverage for the generative features our team validates.
  • Leverage LLM judge scoring output quality in automation, ensuring reliable, repeatable eval jobs that run that produce actionable signal.
  • Validate image/visual generation model output and its associated classification metadata, and detect quality or behavior regressions across model updates.
  • Evaluate natural-language generation to assess whether generated artifacts and responses match user intent, moving at-desk LLM judges into a scalable and repeatable automation environment.
  • Replace exact-match checks for open-ended or factual responses with an LLM-as-judge stage integrated into the pipeline.
  • Assess whether model-generated content (insights and summaries) is sensible and good enough to surface to users.
  • Decide when a component-level check is sufficient and when a full end-to-end user flow is required, and build the tooling for both.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service