We are the Foundation Model Inference Team, within the AI, Search & Knowledge Platform Technologies organization. Our team is responsible for building the inference stack to power Apple Intelligence. It builds frameworks, services, and tools that power Apple's largest foundation models on servers. Our infrastructure powers a wide gamut of services at Apple, including Apple Search, Apple Music, Apple TV, App Store, iMessages, Photos & Camera, Spotlight, Safari, Siri, and upcoming exciting Apple products, serving millions of queries every day with incredible low latencies, drawing every ounce of compute from our hardware. As part of this group, you will get a chance to bring intelligence to billions of users across the world. You will have an opportunity to make a difference in people's lives by empowering them with AI. You will have a chance to work on optimizing billions of parameter language, vision, and speech models using state-of-the-art technologies and make them run at the scale of Apple. Work alongside the Foundation Model Research team to optimize inference for cutting-edge model architectures. Work closely with product teams to build production-grade solutions to launch models serving millions of customers in real-time. Build tools to understand bottlenecks in inference for different hardware and use cases. Mentor and guide engineers in the organization.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior