About The Position

As a Senior/Staff Engineer on the Foundation Model Compute Infrastructure team, you will design and build large-scale infrastructure that powers foundation model training, fine-tuning, evaluation, and inference. You will develop model inference and fine-tuning services, onboard and benchmark new accelerators, and work closely with foundation model researchers and engineers to improve reliability, performance, scalability, and developer productivity across Apple’s AI workloads.

Requirements

  • 5+ years of industry experience building large-scale distributed systems or cloud infrastructure
  • Experience with distributed ML training or inference systems
  • Strong programming skills in Python, Go, C++, or similar systems languages
  • Experience with accelerator infrastructure such as TPU, GPU
  • Experience with Kubernetes, container orchestration, or large-scale cluster management systems
  • Strong communication and collaboration skills across engineering and research teams
  • Bachelor’s degree in Computer Science, Engineering, or related field

Nice To Haves

  • Experience building schedulers, resource managers, or orchestration systems for distributed workloads
  • Familiarity with frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways, or vLLM
  • Experience operating large-scale multi-tenant infrastructure in cloud or hybrid environments
  • Background in performance optimization, fault tolerance, or resource efficiency for large distributed systems
  • Strong expertise in distributed systems, scalability, reliability, and performance engineering
  • Experience designing backend services or infrastructure platforms operating at production scale
  • MS or PhD in Computer Science, Engineering, or related field

Responsibilities

  • Design, build, and evolve large-scale model serving and fine-tuning services for foundation model workloads
  • Develop reliable infrastructure for model deployment, serving autoscaling, traffic management, job execution, container orchestration, and serving performance analysis
  • Improve the performance and usability of model serving and fine-tuning workloads by optimizing latency, throughput, availability, accelerator utilization, checkpoint loading, compilation caching, KV-cache-aware routing, and workflows for launching, monitoring, debugging, evaluating, and deploying models
  • Onboard and Benchmark new accelerator technologies into Apple’s compute infrastructure
  • Collaborate with the Apple Foundation Model team to integrate technologies such as Pathways, Ray, and Beam, or expose them as reliable and scalable services
  • Mentor engineers and partner across teams to influence the technical direction of Apple’s foundation model compute infrastructure
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service