About The Position

NVIDIA’s EDA Infrastructure organization builds and operates the systems that support chip development. We are looking for an engineering manager to lead the team responsible for operational processes and platforms across incident management, maintenance, on-call, issue management, and customer-serving readiness. You will own the roadmap and delivery, from defining how teams work to building the tools they use. You will partner with infrastructure and service owners to improve reliability, reduce manual work, and ensure services are ready to support customers. Your team will use automation, AI, and lessons from operational events to drive improvements.

Requirements

  • BS degree or equivalent experience with 10+ overall years of software engineering or related experience, including 5+ years of engineering leadership managing teams or complex technical programs.
  • Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement.
  • Strong technical judgment in software architecture, platform integration, and engineering tradeoffs.
  • Clear communication with engineers, cross-functional partners, and executive stakeholders.
  • A record of developing engineers, growing teams, and delivering results under pressure.

Nice To Haves

  • Established readiness standards covering service ownership, support coverage, and reliability objectives.
  • Built, integrated, and scaled platforms pertaining the incident, maintenance, customer experience management, along with on-call and production readiness
  • Applied AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation.
  • Supported EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements.

Responsibilities

  • Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results.
  • Set technical direction, prioritize work, and guide execution across engineering and operational disciplines.
  • Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness.
  • Hire and develop engineers and technical leads, building a team with clear ownership and accountability.
  • Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service