About The Position

NVIDIA is looking for a Systems & Software Engineer interested in building and running reliable large scale infrastructure platform services. In this role you will create and support automation that manages infrastructure, as part of our organization running EDA platforms atop of NVIDIA hardware.

Requirements

  • BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics) or equivalent experience
  • Experience coding in Python and/or Go. Note : you must be able to produce code without AI assistance during the interview process.
  • Experience with infrastructure automation and distributed systems design
  • Experience developing tools for running large scale private or public cloud systems in production
  • In depth knowledge of Linux systems and Containers
  • A track record showing a good balance between initiating your own projects, convincing others to collaborate with you and collaborating well on projects initiated by others.
  • 8+ years of relevant experience

Nice To Haves

  • Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive .
  • Experience accelerating positive impact to the business using coding assistant(s), MCP servers, or AI agents
  • In-depth knowledge of networking and storage technologies.
  • Orchestrating existing automated and manual processes using Continuous Integration and/or workflow engines
  • Experience with duplicating or emulating existing complex production infrastructure into staging/test environments.
  • Experience working with or developing bare metal as a service (BMaaS) associated systems
  • Experience working with or developing multi-cloud infrastructure services and running private or public cloud systems based on one or more of Kubernetes, OpenStack, Docker or Slurm .
  • Experience working with Nvidia GPU-based systems

Responsibilities

  • Design, build, deploy, and run infrastructure services & manage the software life cycle in scope to meet our business goals
  • Build and maintain automation to eliminate manual toil and scale output without scaling effort
  • Participate in the definition of our internal facing service level objectives and error budgets as part of our overall observability strategy
  • Practice sustainable blameless incident prevention and incident response while being a member of an on-call rotation
  • Consult with and provide consultation for peer teams on systems design best practices

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service