About The Position

As an SRE on our team, you'll own the reliability, performance, and scale of the distributed storage and data platform systems that power Apple's services. You'll debug replication and consensus failures, tune systems at petabyte scale, and write production code that operates the platform. We firmly believe in ownership, with software engineers accountable for the code they write. The Apple Services Engineering (ASE) organisation builds and provides systems and infrastructure that fuel Apple’s services — iCloud, iTunes, Siri, and Maps. Our team builds and operates the data platform infrastructure behind them, keeping petabyte-scale workloads fast, resilient, and reliable. The platform runs on large-scale distributed systems, including object stores, databases, and data pipelines, on Linux across private and hybrid cloud. You'll work on storage engines, distributed consensus, and data-flow internals, partnering with development teams on system-wide architecture rather than individual components.

Requirements

  • Experience in managing and scaling large-scale distributed systems in a private or hybrid cloud environment.
  • Comfortable designing, writing, and releasing production code in languages such as Go or Python.
  • Able to debug and reason about how distributed systems fail and perform at scale.
  • Willingness to take part in on-call rotations and incident response.
  • A good grasp of Unix internals and networking fundamentals.

Nice To Haves

  • Contributions to distributed-systems internals, open-source data infrastructure, or storage/database engines.
  • Experience defining SLIs/SLOs, building observability, and using error budgets to drive reliability decisions.
  • Experience with data migration, disaster recovery, or capacity planning at scale.

Responsibilities

  • Own the reliability, performance, and scale of distributed storage and data platform systems.
  • Debug replication and consensus failures.
  • Tune systems at petabyte scale.
  • Write production code that operates the platform.
  • Work on storage engines, distributed consensus, and data-flow internals.
  • Partner with development teams on system-wide architecture.
  • Take part in on-call rotations and incident response to keep critical systems healthy.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service