About The Position

This role focuses on contributing to Machine Learning and Generative AI technologies by ensuring the integrity of data powering AI systems at scale. The position involves building and maintaining intelligent systems, validation frameworks, and monitoring pipelines to maintain a healthy data ecosystem, ensuring models are trained, evaluated, and deployed on trustworthy data. The work is foundational to ML features reaching hundreds of millions of users, operating at the intersection of statistical rigor and production systems. Collaboration will occur with ML Engineering, Data Engineering, Privacy, and Legal teams. The role is central to ML and AI quality, encompassing ownership of training and validation datasets, defining and analyzing observability metrics for product insights, and leading telemetry analysis for GenAI workflows to ensure Apple's financial features are built on high-quality data for both conventional ML and generative AI systems. The ideal candidate is detail-oriented, understands that model quality begins with data, possesses strong statistical instincts, recognizes production system issues like silent degradation and data drift, and can translate quality signals into actionable decisions.

Requirements

  • 3+ years of experience in data science or a closely related analytical role, with a strong focus on data quality, model evaluation, or ML observability in production environments.
  • Proficiency in Python (Pandas, NumPy, Scikit-learn) and SQL for complex data analysis, metric creation, and validation.
  • Experience querying and analyzing large-scale datasets using distributed computing frameworks (e.g., PySpark, Spark, or distributed SQL).
  • Solid understanding of statistical methods — hypothesis testing, distribution analysis, data drift detection, and statistical process control.
  • Experience in defining and tracking ML model health metrics in production — model performance monitoring, feature drift detection, and observability instrumentation.
  • Familiarity with GenAI or LLM systems, including common quality failure modes, output evaluation approaches, and telemetry instrumentation.
  • Strong communication skills — ability to translate complex data quality findings and model health risks into clear, actionable insights for both engineering and non-technical stakeholders.

Nice To Haves

  • Experience with data visualization and dashboarding tools (e.g., Tableau, Apache Superset, Databricks) to present complex ML telemetry.
  • Familiarity with LLM evaluation frameworks (e.g. LangSmith) or techniques like LLM-as-a-judge.
  • Experience with Bayesian or causal graph-based approaches to synthetic data generation.
  • Familiarity with confidence calibration techniques and uncertainty quantification.
  • Experience with ML monitoring or observability platforms (e.g., MLflow, Weights & Biases, or equivalent).
  • Experience working with privacy-constrained data or under regulatory compliance frameworks (GDPR, DMA).
  • Background in financial services, fintech, or consumer payment products.

Responsibilities

  • Own the health of the data ecosystem that underpins ML and GenAI features across Wallet, Payments, and Commerce — building validation frameworks, defining observability metrics, and leading telemetry analysis that keeps every model trained, evaluated, and monitored on data teams can trust.
  • Curate, analyze, and maintain gold-standard ground-truth datasets for model evaluation and continuous validation across both ML and GenAI systems.
  • Audit training data for systemic bias and fairness gaps prior to model deployment; establish ongoing analytical checks to catch bias introduced by data drift over time.
  • Define, track, and report key data quality metrics — completeness, accuracy, timeliness, validity — for engineering and leadership audiences.
  • Design and define automated data quality rules and thresholds, partnering with Data Engineering to ensure these checks are integrated into model development and CI/CD workflows.
  • Define and own ML observability metrics — model performance, output distributions, training-serving skew, silent degradation and feature drift — translating raw production signals into actionable insights for engineering and product teams.
  • Design and develop observability dashboards and reporting workflows that give stakeholders a consistent, real-time view of model health across both conventional ML and GenAI systems.
  • Define and analyze telemetry across GenAI workflows, tracking quality signals such as output coherence, latency, task completion rates, and regression patterns.
  • Identify degradation patterns and domain-specific failure modes in GenAI systems through systematic telemetry analysis, translating findings into concrete recommendations for model and data teams.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service