About The Position

The Apple Services Engineering team (ASE) is looking for a Site Reliability Engineer to join the Apple Data Platform SRE team. This team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. The role sits at the intersection of infrastructure, automation, and customer success, involving incident response, hands-on support to internal teams, and partnering with developers to ensure cutting-edge services like Spark, Flink, Airflow, Trino, Notebooks, and LLM-based agent platforms are reliable at scale. This is an opportunity to build deep expertise across a technically diverse platform, specializing in big data engines and catalog/governance layers. The SRE will operate and support the team's full portfolio, from ML/AI platform services to multi-cloud infrastructure, and become a go-to expert for big data platform services. They will also serve as a first point of contact for internal customers, translating technical issues into clear diagnoses and fast resolutions. The ideal candidate is self-motivated, thrives on ownership, enjoys solving operational problems, finds satisfaction in helping customers, and wants to contribute to the scaling of Apple's data engineering platform.

Requirements

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python.
  • Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Excellent written and verbal communication skills, with the ability to explain technical issues clearly to non-expert customers.
  • Solid grounding in SRE principles, with prior on-call, production-support, or customer-facing support role experience.

Nice To Haves

  • Working knowledge of Golang.
  • Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.
  • Prior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.
  • Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
  • Working knowledge of CI/CD pipelines and deployment workflows.
  • Experience with S3 and cloud storage/networking fundamentals.
  • Familiarity with data pipeline orchestration and workflow scheduling patterns.
  • Intellectual curiosity and a drive to keep learning.

Responsibilities

  • Operate and support the team's full portfolio, from ML/AI platform services to multi-cloud infrastructure.
  • Become the team's go-to expert for big data platform services, including Spark, Flink, Airflow, Trino, Notebooks, REST Catalog services, and data governance.
  • Serve as a first point of contact for internal customers, diagnosing and resolving technical issues.
  • Drive the reliability roadmap for a set of services.
  • Automate manual operations through scripting or tooling.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service