Join our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up. During your internship, you may: Build components of our performance analysis and profiling infrastructure. Collect and analyze performance data from our custom ML accelerators, including hardware counters, execution traces, and memory behavior. Develop tooling to trace host-side runtime activity, system behavior, and accelerator execution. Help correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads. Build analysis and visualization tools that help engineers identify performance bottlenecks and optimize models. Work alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and develop tools that improve developer productivity. Representative projects Implement the data collection framework for hardware performance counters on a custom PCIe-based accelerator. Develop a user-space service for low-overhead tracing of accelerator activity. Design and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units. Create an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Intern
Education Level
No Education Listed