Our team builds the ML-inference stack that powers generative AI for Apple Intelligence's Private Cloud Compute — running on Apple Silicon in the datacenter, distributing work across on-SoC acceleration hardware and multi-node clusters. Built on Private Cloud Compute's privacy guarantees, we're growing the team to scale across more platforms and support a widening set of features. As part of the team you will help engineer continuous improvements in stability and performance for Private Cloud Compute, help implement entirely new functionality as it emerges from the research community, and help bring our inference stack up on new generations of SoCs and hardware acceleration IP as we extend to more platforms — in collaboration with hardware, product and research teams throughout Apple. We write performant and scalable frameworks (primarily in Swift, with C++ where we bridge to the hardware) to distribute and coordinate ML inference across the acceleration IP blocks of different SoCs, and to move data and coordinate work reliably across multi-node inference clusters. You will integrate inference code into a full service stack so that user traffic is served reliably and performantly, with a strong focus on code that is easy and safe to develop, update, and monitor in production. We're a collection of highly skilled and friendly engineers who value each other's opinions and experience. We strive for excellence and believe strongly in the quality of our output. We are a team of domain experts, each specializing in specific core subject areas, with broad collective experience across cloud software services and platforms.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior