System Software Engineer, Distributed Systems

NVIDIASanta Clara, CA
$152,000 - $287,500Onsite

About The Position

The VLSI Productivity and Infrastructure team supports over 1000 chip design engineers by building tools and platforms to enhance their daily work. The team's mission is to increase the speed of chip designers. They develop and maintain long-term systems for build automation, observability, analytics, automated error detection and remediation, and codebase modernization, with a strong emphasis on stability. The core workflow infrastructure operates as userspace software on bare-metal Linux hosts, without requiring sudo or containers. State and artifacts are coordinated via NFS, compute-intensive workflows are launched on IBM LSF, and adjacent services for APIs and observability are provided. This is a high-ownership environment where individuals are expected to become experts in their built systems. The role seeks a pragmatic and versatile systems engineer who enjoys low-level work and creating tools for other engineers. It's a generalist position focusing on distributed systems and operational excellence in an environment that operates below the container layer, emphasizing coordination, reliability, performance, and the safe evolution of legacy systems, including the incremental modernization of large codebases into Go. This role involves writing userspace software for managing state, concurrency, and reliability at scale, rather than just configuring CI/CD pipelines.

Requirements

  • B.S. CS/EE (or equivalent experience)
  • 5+ years developing and operating production software in Go and/or Python, ideally in large codebases
  • Strong Linux fundamentals: processes, filesystems, permissions, synchronization/locks, concurrency, and debugging
  • Solid distributed-systems thinking: failures, retries/timeouts, backoff, idempotency, and operational rigor
  • Experience building long-runtime automation or services on shared compute clusters (batch schedulers, build systems)
  • Ability to translate ambitious, high-level goals into a safe delivery plan (instrumentation, staged rollout, measurable outcomes)

Nice To Haves

  • Hands-on experience with shared filesystems at scale (NFS), or coordination patterns on eventually-consistent storage
  • Experience with batch job scheduling, shared compute fleets, or build systems
  • Track record of incremental modernization (tests, shadow runs, canaries, rollback plans)
  • Experience partitioning/optimizing metadata-heavy systems and reducing I/O or R/W hot spots
  • Strong incident/debug tactics: clear root-cause analysis, remediation, and guardrails as well as rapid comprehension and ownership of unfamiliar codebases in any language (including LLM-generated code) to implement high-leverage changes

Responsibilities

  • Design, build, and deliver core components of our next-generation productivity platforms
  • Develop reliable userspace infrastructure for long-running engineering workflows at scale on bare-metal Linux hosts
  • Build state coordination over NFS (atomicity, idempotency/dedup, partial-write recovery, without privileged ops)
  • Build and improve orchestration around IBM LSF (submission/tracking, retries/cancel, log capture, fairness/backpressure)
  • Convert legacy codebases into modern powerhouses using incremental migration techniques (e.g., Perl to Go), with stage gates, parity strategies, and strong observability
  • Debug and improve performance and reliability across Linux and Kubernetes, including operational tooling
  • Collaborate with engineering users to turn ambiguous workflows into durable production systems

Benefits

  • Competitive salaries
  • Generous benefits package
  • Equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service