About The Position

Our Scene Understanding team strives to turn cutting edge research into compelling user experiences that realize the potential of generative models to transform creative workflows and smart assistants, fundamentally shifting how people interact with devices and communicate. We are looking for senior technical leaders experienced in architecting and deploying production scale multimodal ML. An ideal candidate has the ability to lead diverse cross functional efforts ranging from ML modeling, prototyping, validation and private learning. Solid ML fundamentals and an ability to place research contributions with respect to state of the art would be an essential part of the role. Experience with training and adapting large language models would be an important need. We are the Intelligence System Experience (ISE) team within Apple’s software organization. The team works at the intersection between multimodal machine learning and system experiences. For example, experiences like Spotlight Search, Photos Memories, Generative Playgrounds, Stickers, Smart wallpapers, etc are all areas that the team has had a significant part in delivering through ML core technologies. These experiences that our users enjoy are backed by production ML workflows, which our team works to scale through distributed training. Additionally, our team also focuses on approaches to optimizing and adapting LLMs to best suit on-device user experiences.

Requirements

  • M.S. or PhD in Computer Science or a related field such as Electrical Engineering, Robotics, Statistics, Applied Mathematics, or equivalent experience.
  • Hands on experience training LLMs/adapting pre-trained LLMs for downstream tasks & alignment.
  • Modeling experience at the intersection of NLP and vision.
  • Proficiency in ML toolkit of choice, e.g., PyTorch.
  • Strong programming skills in Python.

Nice To Haves

  • Familiarity with distributed training.
  • Strong programming skills in C/C++ or ObjC.

Responsibilities

  • Training large scale multimodal (2D/3D vision-language) models on distributed backends.
  • Deployment of compact neural architectures efficiently on device.
  • Learning policies that can be personalized to the user in a privacy preserving manner.
  • Ensuring quality in the wild, with an emphasis on fairness and model robustness.
  • Enriching multimodal capabilities of large language models.
  • Aligning image/video content to the space of LMs for visual actions & multi-turn interactions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service