About The Position

StackYak is building the infrastructure layer for AI, creating software that integrates compute, GPU infrastructure, networking, and inference into a single product. This is an early-stage, funded company focused on real production workloads from the outset. This role is a founding position within its discipline, focusing on the systems layer beneath the product. The successful candidate will have significant autonomy in shaping this layer. The company values individuals who can quickly grasp new concepts, operate effectively without perfect requirements, and take initiative. The role requires an individual who can build the systems that bridge GPU capacity and production infrastructure. This involves managing diverse compute resources, including bare metal hardware (with its associated firmware, drivers, and host management) and cloud-based capacity (where control is limited to software). The core challenge lies in abstracting the complexities between these different forms of capacity, determining what needs to remain visible, and defining how much the rest of the company needs to interact with the infrastructure. The ideal candidate will be proficient in both low-level systems (Linux kernel, firmware, drivers, virtualization, host hardening) and cloud provider APIs, with a strong emphasis on automation at scale. Workloads will span bare metal, virtual machines, and containers, requiring the ability to select the appropriate environment for each. This is a hands-on role involving system design, implementation, debugging, and operation, not solely an architectural position.

Requirements

  • Production Linux experience at a deep level, where solutions are found in the kernel.
  • Experience operating GPU infrastructure for demanding workloads (AI, HPC, cloud).
  • Experience managing NVIDIA and/or AMD platforms through challenging driver or firmware upgrades.
  • Experience provisioning bare metal at scale where manual configuration is not feasible.
  • Experience managing and recovering remote machines that cannot be physically accessed.
  • Experience choosing between bare metal, KVM, and containers with a clear rationale.
  • Experience writing production-level Terraform and Python for tooling, not just simple scripts.
  • Experience hardening hosts and isolating workloads for security-sensitive customers.
  • Experience debugging complex failures involving hardware, firmware, OS, drivers, network, or workload.
  • Experience making infrastructure decisions with incomplete information and evaluating their outcomes over time.

Nice To Haves

  • Experience building or operating infrastructure at a neocloud, hyperscaler, GPU cloud, HPC environment, hosting company, data-center operator, or AI infrastructure company.
  • Experience running a fleet that mixes owned hardware with rented capacity and understanding the cost implications.
  • Experience with multi-GPU and multi-node systems.
  • Understanding of NUMA, PCIe topology, RDMA, NIC placement, and their impact on AI workloads.
  • Experience specifying or accepting hardware and understanding the process from purchase order to a bootable machine.
  • Experience designing infrastructure for multi-customer safety.
  • Experience working outside of tightly siloed enterprise teams.
  • Personal projects demonstrating a passion for systems.

Responsibilities

  • Define and manage the compute fleet, including owned versus rented resources, standardization, and cost implications.
  • Manage the machines themselves, including Linux, kernel, firmware, drivers, and GPU software (CUDA/ROCm) at a deep level.
  • Own the boundary between bare metal, VM, and container environments, adapting solutions as needed.
  • Ensure workload and customer isolation, validating security claims against potential attacks.
  • Develop automated provisioning and recovery systems for machines, minimizing manual involvement.
  • Integrate with cloud providers and other compute vendors, understanding their offerings versus advertised capabilities.
  • Oversee the process of specifying hardware and bringing it online through vendors and data-center partners.
  • Troubleshoot and resolve failures that span multiple layers, including hardware, firmware, OS, virtualization, storage, and network.
  • Make the infrastructure operable by others, preventing the fleet from becoming a collection of undocumented, one-off systems.

Benefits

  • Competitive compensation
  • Meaningful equity
  • Direct access to founders
  • Remote and distributed work environment
  • Opportunity to shape product infrastructure
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service