We are seeking a highly skilled Staff Failure Analysis Engineer to lead complex investigations of server and datacenter hardware failures, driving root cause identification and corrective actions that improve product quality and reliability. In this role, you will work hands-on with server systems and components including CPUs, GPUs, DIMMs, NVMe drives, FPGA, NICs, and power supplies, performing troubleshooting, reliability testing, data analysis, and debug activities throughout the product lifecycle. You will partner closely with engineering, manufacturing, quality, and test teams to resolve issues, enhance manufacturing processes, support new product introductions, and provide technical leadership in a fast-paced, high-volume manufacturing environment. The ideal candidate brings strong server hardware expertise, failure analysis experience, and a passion for delivering reliable, high-performance technology solutions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior