About The Position

Apple Service Engineering (ASE)'s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies — ensuring platform availability, reliability, and performance at Apple scale. You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains.

Requirements

  • 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems.
  • Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes.
  • Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale.
  • Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt).
  • Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering.
  • Excellent verbal and written communication skills with the ability to influence across teams and levels.
  • Demonstrated ability to recruit, develop, and retain high-performing engineering talent.

Nice To Haves

  • Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation.
  • Experience managing or scaling batch compute, job scheduling, or HPC platforms.
  • Proficiency in Go or Python with a strong automation-first mindset.
  • Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale.
  • Experience operating large-scale multi-tenant Infrastructure as a Managed Service.
  • Experience managing geographically distributed teams and follow-the-sun on-call models.
  • Track record of driving capacity efficiency initiatives resulting in measurable cost optimization.

Responsibilities

  • Operate and scale large-scale batch compute infrastructure across Apple's data centers.
  • Manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies.
  • Ensure platform availability, reliability, and performance at Apple scale.
  • Drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization.
  • Balance operational excellence with engineering innovation.
  • Establish SLOs and drive production readiness.
  • Build automation and tooling to enable a growing platform to scale efficiently.
  • Champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service