About The Position

The On-Device Machine Learning team at Apple transforms groundbreaking research into practical applications, enabling billions of Apple devices to run powerful AI models locally, privately, and efficiently. This team operates at the intersection of research, software engineering, hardware engineering, and product development. They build essential infrastructure for machine learning at scale on Apple devices, including onboarding innovative architectures to embedded systems, developing optimization toolkits for model compression and acceleration, building ML compilers and runtimes for efficient execution, and creating comprehensive benchmarking and debugging toolchains. This infrastructure supports Apple’s machine learning workflows across Camera, Siri, Health, Vision, and other core experiences, contributing to the Apple Intelligence ecosystem. The role is for an ML Infrastructure Engineer with a focus on model compilation, working closely with model authoring, runtime, and performance teams to ensure models can leverage the full capabilities of the hardware. The team is building an end-to-end developer experience for machine learning development using Apple’s vertical integration, covering model authoring, optimization, transformation, execution, debugging, profiling, and analysis. This specific role focuses on the core runtime for execution across various devices and use cases. The ideal candidate is a creative, versatile, and passionate software engineer interested in machine learning, common compiler optimizations, and system software engineering. The team utilizes an MLIR-based compiler stack to target the neural engine, GPU, and CPU for ML workflows and execution.

Requirements

  • 3-5 years working on MLIR-based compilers.
  • Familiarity with common ML model architectures, execution schemes, and operations.
  • Familiarity with C++
  • Familiarity with PyTorch or related training frameworks

Nice To Haves

  • Familiarity with Swift.
  • Familiarity with programming paradigms for the GPU, CPU, and Neural Engine.
  • Familiarity with writing kernels for ML model execution.

Responsibilities

  • Work closely with model authoring, runtime, and performance teams to ensure models can bring to bear the full capabilities of the hardware.
  • Focus on the core runtime for execution across a wide variety of devices and use cases.
  • Utilize an MLIR-based compiler stack to target the neural engine, GPU, and CPU for ML workflows and execution.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service