About The Position

The Apple Services Engineering (ASE) team is responsible for powering services like the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books at a massive scale. Within ASE, the Apple Data Platform SRE team maintains a large, multi-cloud platform used by thousands of internal engineers for data and AI product development. This role is at the intersection of infrastructure, automation, and customer support, involving incident response, providing hands-on assistance to internal teams, and collaborating with developers to ensure the reliability of services such as Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms across AWS, GCP, and on-premise Kubernetes.

Requirements

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python.
  • Strong hands-on AWS experience: IAM (roles, policies, permission boundaries, KMS), EKS, RDS, S3, VPC networking/endpoints, autoscaling groups, EBS.
  • Kubernetes administration experience — RBAC, node/pod scheduling, autoscalers, PriorityClasses/PDBs, and troubleshooting cluster-wide disruptions.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles.
  • Prior on-call or production-support experience.

Nice To Haves

  • Working knowledge of Golang.
  • Experience with Infrastructure-as-Code (Crossplane and/or Terraform), including debugging state drift and composition/controller issues.
  • Experience with GitOps workflows (Flux or similar) — HelmRepository/reconciliation troubleshooting and Helm chart deployment.
  • Multi-cloud exposure (GCP) — parity and migration scenarios.
  • Experience with Splunk for log pipeline debugging (e.g., fluent-bit).
  • Familiarity with Spark/Flink running on Kubernetes (executor scheduling, node affinity).
  • Comfort with GitHub PR review workflows in an infrastructure-as-code / GitOps context.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning.

Responsibilities

  • Operate and support the team's full portfolio, from big data pipelines to ML/AI platform services.
  • Become the team's go-to expert for multi-cloud infrastructure, including AWS core services (IAM, EKS, RDS, S3, VPC networking, autoscaling, EBS), Kubernetes administration at scale, and Infrastructure-as-Code and GitOps tooling.
  • Respond to incidents, troubleshoot issues with IAM policies, cluster scheduling, and GitOps reconciliation.
  • Drive initiatives to reduce failure modes over time.
  • Take ownership of services, develop their reliability roadmaps, and collaborate with the team on broader strategic directions.
  • Deep dive into cloud and Kubernetes internals.
  • Act as a trusted expert for internal customers.
  • Contribute to the evolution of Apple Data Platform's multi-cloud infrastructure.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service