About The Position

At Apple, great ideas turn into phenomenal products, services, and customer experiences at a pace few companies can match. We're looking for a seasoned engineering leader to build and lead a team of distributed systems engineers creating a brand-new service that will power experiences for Apple customers at massive scale. In this role, you will set the technical and organizational direction for the project from the ground up. You'll partner closely with the leaders of Search, Client, AI and Data Engineering teams to define the architecture, roadmap, and execution plan for a mission-critical service. You will be accountable for architecting, building, and operating highly distributed systems and microservices that are low-latency, fault-tolerant, deeply observable, and backed by strongly consistent systems of record. You'll also lead the effort to bring ML workloads into production, making sure models are served reliably and efficiently in the request path. This is a senior leadership role for someone who thrives on ambiguity, enjoys building teams from scratch as much as systems, and can grow a high-performing organization while staying hands-on with architecture and technical decision-making.

Requirements

  • Bachelor's or Master's degree in Computer Science or a related field, or equivalent experience.
  • 15+ years of professional software development experience, including building and operating large-scale distributed systems in production.
  • 8+ years of people management experience leading teams of senior backend or distributed systems engineers.
  • Proven track record of taking a new service from concept to production at scale and owning its long-term operation.
  • Deep expertise in architecting multi-tiered distributed systems and microservices: API design, authentication and authorization, concurrency, scaling for high availability, fault tolerance, and reliability.
  • Strong background in the Java/JVM ecosystem (e.g., Spring Boot, async/reactive frameworks such as Netty or Project Reactor) and RPC frameworks such as gRPC.
  • Deep understanding of transactional consistency models, ACID semantics, and the trade-offs between relational and NoSQL database technologies.
  • Experience with event streaming and queueing systems (e.g., Kafka), stream processing, and caching technologies (e.g., Redis) in production services.
  • Experience running services on AWS or GCP with cloud-native tooling (Docker, Kubernetes) and mature CI/CD pipelines.
  • Strong operational mindset: experience defining SLOs, building observability (metrics, logging, tracing, alerting), and leading incident response for customer-facing services.
  • Experience deploying and operating ML models or ML-powered services in production environments.
  • Excellent communication skills and the ability to influence and align senior cross-functional partners.

Nice To Haves

  • Experience building search, ranking, recommendation, or personalization systems.
  • Hands-on experience serving and optimizing LLMs or ML models in the request path, including GPU-backed inference at scale.
  • Familiarity with MLOps practices: model versioning, feature stores, A/B testing and experimentation, and model monitoring.
  • Experience partnering with client and mobile teams on API design, data contracts, and end-to-end performance.
  • Experience with data lake and analytics technologies (e.g., Iceberg, Spark) and schema tooling (Protobuf, Avro, Schema Registry).
  • Background in security and privacy: TLS, X.509 certificates, OAuth2/OIDC, threat modeling, and privacy-preserving system design.
  • Experience growing a team from early stage to a multi-team organization, including developing senior engineers and future leaders.
  • Self-motivated and comfortable with ambiguity, with experience in fast-paced, agile environments.

Responsibilities

  • Set the technical and organizational direction for the project from the ground up.
  • Partner closely with leaders of Search, Client, AI and Data Engineering teams to define the architecture, roadmap, and execution plan for a mission-critical service.
  • Architect, build, and operate highly distributed systems and microservices that are low-latency, fault-tolerant, deeply observable, and backed by strongly consistent systems of record.
  • Lead the effort to bring ML workloads into production, ensuring models are served reliably and efficiently in the request path.
  • Build and lead a team of distributed systems engineers.
  • Grow a high-performing organization.
  • Stay hands-on with architecture and technical decision-making.
  • Hire and grow teams.
  • Take a new service from concept to production at scale and own its long-term operation.
  • Define SLOs, build observability (metrics, logging, tracing, alerting), and lead incident response for customer-facing services.
  • Develop senior engineers and future leaders.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service