About The Position

Apple's Responsible AI and Safety team focuses on innovative technologies, methodologies, and research to enable fantastic user experiences and to push the frontier of machine learning. Our team is looking to hire a leader with a strong track record in Applied Research, who is passionate about ML and foundation models with a focus on responsibility, fairness, and safety. In this role, you will lead the research and application of ML methods for technologies that power breakthrough user experiences while upholding Apple's values, privacy, and quality standards. This role leads Apple's Quality Platform Engineering function within Responsible AI — the team responsible for giving engineering teams across Apple Intelligence fast, trustworthy safety signal before code merges, and for extending safety testing coverage to every hardware platform and form factor Apple ships. Most of this role's impact will come from two things: building a strong team and a healthy team culture from the ground up, and establishing the flywheel that connects Quality Platform Engineering to Product Evaluations & Research and Post-Ship Insights — so fast signal, coverage gaps, and pipeline improvements consistently translate into real engineering decisions and clear leadership visibility, rather than one-off dashboards nobody acts on. You should be technically fluent — comfortable with evaluation pipelines, production ML and agentic systems, and cross-platform testing — enough to earn credibility with the team, ask sharp questions, and evaluate hard tradeoffs with good judgment. This role moves at a fast pace, and you should be comfortable making product recommendations in ambiguous situations, often with limited or imperfect data. Just as important are excellent communication, strong product sense to prioritize a team's limited capacity against Apple's highest-risk platforms and use cases, and a track record of hiring and developing technical teams.

Requirements

  • 5+ years of technical team management or leadership experience
  • Experience with ML evaluation, production ML systems, or test/CI infrastructure at scale
  • Strong engineering skills and experience writing production-quality code (Python or similar)
  • Experience working across multiple platforms or hardware form factors, or a demonstrated ability to ramp quickly across unfamiliar platforms
  • Experience working with human-labeled or crowd-sourced evaluation data, including reasoning about label noise and inter-rater agreement
  • Experience working on Responsible AI, AI safety, or trust & safety-adjacent engineering
  • Experience with generative model evaluation and common failure modes
  • Strong organizational and operational skills working with large, multi-functional, diverse teams
  • MS or PhD in Computer Science, Machine Learning, Statistics, or related field, or equivalent experience

Nice To Haves

  • Prior exposure to ML evaluation, test infrastructure, or safety-adjacent engineering work is strongly preferred.
  • Familiarity with hardware/platform-specific testing considerations (e.g., on-device constraints, new form factors)

Responsibilities

  • Hire, grow, and build the culture of a team spanning test engineering, applied ML, and platform-specific expertise
  • Own leadership reporting and safety sign-off representation for Quality Platform Engineering - translating pipeline health, coverage, and findings into risk and product terms for non-technical stakeholders
  • Set and continuously reprioritize which feature clusters, platforms, and hardware form factors the team covers, based on risk, launch timing, and known coverage gaps
  • Drive investment in test pipeline efficiency — turnaround time, infra cost, and flakiness — so pre-merge signal stays fast enough to be useful in the dev loop
  • Extend safety testing coverage beyond iPhone/iOS to every platform and device category Apple ships, partnering with platform teams to close hardware-specific gaps
  • Upstream trusted datasets for product-critical safety cases, so component teams have reliable ground truth to test against early
  • Champion shift-left safety testing — catching issues at the component level before they reach end-to-end evaluation
  • Stay technically fluent enough to evaluate findings, unblock the team on hard problems, and track how model and pipeline behavior shifts across releases
  • Ensure the team's signal stays consistent with Product Evaluations & Research's ship-time metrics, so pre-merge results reliably predict end-to-end outcomes rather than drifting from them
  • Manage occasional exposure to sensitive or policy-relevant content surfaced during testing, and support team wellbeing around that exposure
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service