Principal Firmware Engineer - Data Center Server Management

NVIDIASanta Clara, CA
$272,000 - $488,750

About The Position

NVIDIA is seeking an expert engineer to design rack-level solutions for next-generation scaling AI supercomputing platforms. The role involves owning the end-to-end manageability architecture for these products in data centers. You will collaborate with internal and external component leads, drive customer use cases, align architecture with customer requirements, and ensure the delivery of high-quality products to market. This is an opportunity to join NVIDIA at the forefront of technological advancement in AI computing.

Requirements

  • 15+ years of relevant experience working on server firmware (BMC) and platform software development.
  • BS, MS, or PhD in EE/CS or related field of education or equivalent experience.
  • Hands-on experience with data center health management workflows.
  • Proven record of delivering server firmware for large data centers.
  • Strong knowledge of data center management, server architecture, and server manageability in data centers.
  • Strong and demonstrable skill in C/C++ and Python.
  • Experience with programming and debugging skills for server platforms.
  • Experience in SCM (e.g. Git, Perforce) and project management tools like Jira.
  • Excellent written and oral communication skills, good work ethics, high sense of team-work, commitment to produce quality work and finish tasks daily.
  • Self-starter who loves to find creative solutions to complicated problems and is hands-on with coding.

Nice To Haves

  • Hands-on experience with data center health management.
  • Hands-on with x86 or ARM system architecture.
  • Proven technical leadership to drive large complex problems with 50+ engineers working.

Responsibilities

  • Drive server management for large clusters and data centers deploying GPUs and Grace solutions from Nvidia.
  • Work with data center architects and cloud customers to narrow down requirements for implementation to ensure speed of light product development.
  • Work with internal teams to ensure requirements are designed and implemented correctly in each firmware and software module.
  • Collaborate with other leads to design & build data center health management workflows.
  • Drive reliability and optimization in firmware architecture from a data center viewpoint.
  • Work closely with the cluster bring-up team and resolve issues at Speed of Light.
  • Own firmware delivered to data centers in terms of quality, reliability, and telemetry performance.

Benefits

  • Equity
  • Benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service