StackYak is building the infrastructure layer for AI, creating software that integrates compute, GPU infrastructure, networking, and inference into a single product. This is an early-stage, funded company focused on real production workloads from the outset. This role is a founding position within its discipline, focusing on the systems layer beneath the product. The successful candidate will have significant autonomy in shaping this layer. The company values individuals who can quickly grasp new concepts, operate effectively without perfect requirements, and take initiative. The role requires an individual who can build the systems that bridge GPU capacity and production infrastructure. This involves managing diverse compute resources, including bare metal hardware (with its associated firmware, drivers, and host management) and cloud-based capacity (where control is limited to software). The core challenge lies in abstracting the complexities between these different forms of capacity, determining what needs to remain visible, and defining how much the rest of the company needs to interact with the infrastructure. The ideal candidate will be proficient in both low-level systems (Linux kernel, firmware, drivers, virtualization, host hardening) and cloud provider APIs, with a strong emphasis on automation at scale. Workloads will span bare metal, virtual machines, and containers, requiring the ability to select the appropriate environment for each. This is a hands-on role involving system design, implementation, debugging, and operation, not solely an architectural position.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed