About The Position

The Site Reliability Engineering team within Apple Ads ensures the reliability, performance, and availability of ML Platform and Services at scale. The team partners closely with Ads engineering, data science and ML platform teams to enable product delivery through design, configuration, and automation of machine learning infrastructure powering Apple Ads applications. We are looking for a ML Platform Infrastructure Engineer to help build and evolve the next generation of Apple Ads machine learning platform — enabling fast, reliable, and scalable operations across AWS-based environments supporting transactional and analytical workloads.

Requirements

  • 3+ years of experience in internet-facing backend production systems, SRE or ML Operations focused roles on large scale distributed cloud infrastructure
  • Proven expertise with AWS-managed infrastructure
  • Familiarity with ML lifecycle and associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc.
  • Strong programming skills in at least one of: Python, Java, Rust, Go or similar languages
  • Hands-on experience with Linux systems and deep knowledge of its internals.
  • Demonstrated experience with Infrastructure as Code, especially Terraform.
  • Strong foundation in SRE concepts: Monitoring, alerting, observability, Incident response and root cause analysis, Error budgets, SLAs/SLOs, and system reliability

Nice To Haves

  • Built tools or services that automate platform operations, reduce toil, or improve cost efficiency.
  • Experience managing Kubernetes clusters at scale in production environments.
  • Hands-on experience troubleshooting distributed systems under real-world load.
  • Clear communication skills and comfort collaborating across engineering, infrastructure, and product teams.
  • AWS certifications or broad experience across multiple AWS services is a plus.
  • Understanding of modern GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia ), associated drivers
  • Understanding of high-performance fabrics and network architecture, power, and thermal limits

Responsibilities

  • Own the health, performance, and scalability of large scale infrastructure powering ML training, inference, serving workloads and associated platform tooling.
  • Build automation that eliminates manual processes, improves platform resilience, and enables teams to move faster with confidence.
  • Build platform solutions, not just configure pipelines.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service