The Accelerator Reference Design Team is looking for a System Engineer to design, implement, and support hardware management for custom AI hardware. The ARD team is working on the design and implementation of hardware modules for Meta's custom AI accelerators. Meta is developing large-scale AI/HPC clusters using custom-designed AI accelerators. In this role, you will have a unique opportunity to shape the future AI/HPC of Meta by specifying technical requirements for device management, driving specifications and designs for integrating device management into fleet management, and steering the industry and ecosystem partner direction. The ideal candidate will work in a cross-functional engineering environment, prioritizing competing workstreams based on impact, deadlines, and stakeholder needs. They will have experience in solving complex device management problems in high-performance computing that span across silicon, hardware, and software. They will also have direct experience in hardware design and management of complex ASICs. A successful candidate will have experience working on hardware management across rack, server, and ASIC levels. The candidate will have experience with current industry practices and standards for device management and debug. The position requires a lead developer to architect solutions, debug complex issues, define scope for open-ended technical challenges, and deliver hardware management at scale. The candidate will work closely with cross-functional stakeholders and partners who are on the front-line of developing Meta's custom ASICs for AI. The Accelerator Reference Design Team designs, builds, brings-up, tests and integrates hardware systems that power Meta's custom AI silicon platforms, deployed in data centers worldwide. In this role, you will help design and build open and efficient AI platforms deployed at scale.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior