Senior System Integration and Validation Engineer

NVIDIA•Santa Clara, CA
•$168,000 - $310,500•Hybrid

About The Position

The Silicon Co-Design Group (SCG) operates at the intersection of architecture, design, marketing, operations, and productization, influencing early architecture through final product delivery across various markets including Datacenter, Gaming, Robotics, Automotive, and Embedded. System Integration is a critical layer where architectural, design, software, and manufacturing decisions converge. This role is crucial for leading system validation, debug, and cross-functional alignment on significant silicon programs, directly impacting the success of flagship builds. The position focuses on two key challenges: proactively identifying and resolving critical silicon and platform issues before customer delivery by enhancing validation coverage through system stress, PVT, and feature-interaction tests, and building trustworthy AI-enabled validation capabilities for debug, triage, root-cause hypothesis generation, and regression workflows that can inform production decisions.

Requirements

  • BTech/BE or MTech/ME in Electronics, Electrical, or Computer Engineering (or equivalent experience).
  • 8+ years of post-silicon validation, system integration, or platform debug experience on shipped GPU, CPU, or SoC products.
  • Hands-on debug experience across silicon, board, and software boundaries, including logic design, signal integrity, power delivery, high-speed I/O, and PVT behavior.
  • Strong EE fundamentals in SI/PI, power delivery, and thermal management.
  • Working understanding of GPU, CPU, or SoC architecture in PC, Datacenter, or Automotive contexts.
  • At least one specific example of resolving an ambiguous, multi-team failure to root-cause closure with a productized fix.
  • Track record of leading a project team end-to-end through a crisis with owned decision points.
  • Demonstrated experience building or scaling AI-driven validation workflows (beyond just using them) with adoption beyond oneself and measurable impact on debug velocity, coverage, or escape rate.
  • Ability to describe AI workflow guardrails and identify where AI is dangerous in the workflow.

Nice To Haves

  • History of building reusable validation methodology, debug playbooks, or test frameworks adopted by other programs or teams (supported by patents, conference papers, invited talks, or recognized contributions).
  • Hands-on subsystem depth in HBM, SerDes or high-speed I/O, power and thermal, or advanced packaging, including failure modes, debug instrumentation, and production tradeoffs.
  • Experience partnering deeply with a counterpart team in India or another major engineering hub, with examples of shared culture, metrics, and on-call responsibilities across geographies.
  • AI work beyond personal-copilot use, such as agentic workflows, RAG-grounded debug assistants, regression bucketing, or automated triage deployed at team scope with adoption metrics.

Responsibilities

  • Own end-to-end system validation of NVIDIA GPU, SoC, and platform programs, including feature checks, PVT stress, large-scale system-stress campaigns, and multi-unit fleet testing.
  • Lead debug efforts for complex cross-stack issues (logic, signal integrity, power delivery, firmware, software interaction) and drive them to root cause with reusable workarounds and productized fixes.
  • Develop test plans, scripts, and automation for next-generation chips prior to physical builds, translating architectural requirements, boot flows, high-speed I/O, and power/thermal dependencies into executable coverage.
  • Design and deploy AI-enabled debug and validation workflows with explicit guardrails (evals, regression validation, false-positive handling) to improve cycle time, debug velocity, or escape rate.
  • Translate complex silicon and system risks into clear, decision-ready options for technical and executive audiences.
  • Collaborate with globally distributed teams across silicon design, DFT, firmware, software, manufacturing, and platform engineering, including counterpart engineers in India and other major hubs.
  • Lead crisis-mode debug task forces for stuck programs and institutionalize learnings into reusable methodologies for adjacent programs.

Benefits

  • Competitive benefits
  • Flexible time off
  • Continuous learning
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service