About The Position

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in an area that's shaping the future of how Apple builds and operates AI. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from big data pipelines to multi-cloud infrastructure, and grow into the team's go-to expert for ML/AI platform services — including Ray training and serving, LangGraph agent deployments, RAG architectures, embeddings platforms, and vector store platforms. You won't be building the models yourself, but you'll be the infrastructure backbone behind the teams who do — keeping their services, pipelines, and platforms running flawlessly in production so they can focus on innovation. We're looking for a self-motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love solving hard operational problems, enjoy being the trusted expert customers turn to, and want a front-row seat to Apple's ML/AI infrastructure evolution, this role offers real room to grow your scope and impact over time.

Requirements

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python.
  • Experience with Kubernetes.
  • Experience with at least one major cloud provider (AWS or GCP).
  • Exposure to operating or supporting ML pipelines, model-serving infrastructure, or LLM-based systems in production.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on-call or production-support experience.

Nice To Haves

  • Working knowledge of Golang.
  • Hands-on experience operating or supporting Ray (training/serving), LangGraph or similar agent orchestration frameworks, RAG architectures, embeddings platforms, or vector store platforms.
  • Familiarity with MCP-based tooling and ML lifecycle/dataset management systems.
  • Experience with S3 and cloud storage/networking fundamentals.
  • Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
  • Working knowledge of CI/CD pipelines and deployment workflows.
  • Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.

Responsibilities

  • Operate and support the team's full portfolio, from big data pipelines to multi-cloud infrastructure.
  • Grow into the team's go-to expert for ML/AI platform services — including Ray training and serving, LangGraph agent deployments, RAG architectures, embeddings platforms, and vector store platforms.
  • Keep services, pipelines, and platforms running flawlessly in production.
  • Solve hard operational problems.
  • Be the trusted expert customers turn to.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service