System Engineer Datacenter GPU

Ampcus Inc.Austin, TX
Onsite

About The Position

Ampcus Inc. is seeking a highly motivated System Engineer Datacenter GPU to join their talented team. The role is within IPP (Infrastructure, Planning and Process) Sanity Engineering, a core software infrastructure organization that supports various other groups within Client Software, including Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence, and Driverless Cars. This group provides cloud services that support nearly half a million automated jobs daily on thousands of servers, enhancing the productivity of thousands of Client software engineers globally. The cloud environment is heterogeneous, featuring a mix of machines and devices with diverse operating systems (Windows/Linux/Android) and hardware platforms, including Client GPUs and Tegra Processors. The ideal candidate is passionate about distributed infrastructure, seeks a sophisticated workspace, and is ready to build the next generation of cloud services for chip bringups, design creative solutions, and identify and resolve problems through data analysis.

Requirements

  • Bachelor's or Master's Degree in Computer Science or Software Engineering, or equivalent demonstrable experience.
  • 10+ years of relevant experience.
  • Ability to analyze and debug source code to triage, root cause and resolve issues in the infrastructure.
  • Collaborate with the development teams in improving the build and test infrastructure.
  • Familiar with maintenance and setup of Linux, Windows hosts and popular open source applications such as Nginx, Apache HTTP, Apache Tomcat and MySQL server.
  • Hands-on programming experience with any including but not limited to Python (preferred), JAVA etc.
  • Unix & TCL shell proficiency is expected.
  • Experience in MySQL/No-SQL (plus), should be able to write complex queries.
  • Experience with version control systems like Perforce, GIT.
  • Demonstrable experience working in large scale enterprise production systems.
  • 5+ years of operational experience required.

Nice To Haves

  • Background with automating bare metal and VM provisioning.
  • Prior knowledge of VM isolation for GPUs and Client Confidential Computing is a plus.
  • Experience with public clouds (AWS, GCP, Azure), VM and container virtualization technologies like VMware, KVM, Docker and Kubernetes Clusters.
  • Experience with debugging GPU performance issues, embedded device software development and automation, software driver development and CUDA/TensorRT applications.

Responsibilities

  • Develop framework and scripts to automate workflows and deployments in the private cloud environment.
  • Enhance automation for farm-wide updates, with a thorough understanding of Client GPU hardware and driver stack, SBIOS, and VBIOS.
  • Solve complex problems involving multi-site distributed infrastructure scaling.
  • Lead GPU product bringups (PCIe & Enterprise) in infrastructure.
  • Integrate GPU test suites to infrastructure harness.
  • Deploy and maintain a large farm of machines using Configuration Management & Infrastructure Automation (IaC) tools like Chef, Ansible, Terraform.
  • Develop extensive monitoring dashboards systems for real-time monitoring of infrastructure subsystems using telemetry.
  • Lead a service charter completely, responsible for the development, monitoring, and automation of that infrastructure.
  • Automate and performance tune regression test frameworks.
  • Create self-healing/automated recovery solutions for multi-geo regression farms.
  • Assist in the roll-out and deployment of new development features to support the latest Client hardware and technologies.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service