NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. We are the Datacenter System Software team, and we are looking for a highly motivated, creative Senior Engineer to drive Fleet Scale Debuggability end to end. You will design, architect, and build infrastructure, tooling, analytics on how to collect multi-rack scale logs. The solution should normalize, correlate, and reason over logs spanning multiple components, trays, or racks including NVIDIA's GPUs, CPUs, Network products. The logs shall be fetched inband or out of band and should help triage fleet level issues seen by our customers. Your work directly shortens the path from a raw, noisy log stream to an actionable root cause. Join us at the forefront of technological advancement.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Ph.D. or professional degree