StackYak is building the infrastructure layer for AI. We are an early-stage, funded company building software that brings compute, GPU infrastructure, networking, and inference together into one product. The opportunity is large, the market is moving quickly, and we are building for real production workloads from the start. This is not internal IT. This is not a slow-moving infrastructure team maintaining someone else's platform. The infrastructure is the product. We are a small, senior team with very little bureaucracy. This is a founding role in its discipline. You will be the first person here whose primary responsibility is the systems layer beneath the product, and the shape it takes will largely be yours to decide. Treat this document as a starting point rather than a boundary. The people who do well here take ground early and are not asked to give it back. We move quickly. We do not have months for someone to learn the fundamentals of their discipline. You should already be very good at what you do, be able to ramp into adjacent areas quickly, and be comfortable operating without perfect requirements or neatly defined boundaries. We need someone who can take a model, a set of GPUs, and a production requirement and determine how that model should actually run. Hand you a frontier open-weight model — dense or mixture-of-experts, tens to hundreds of billions of parameters — along with the hardware actually available and the latency a customer actually needs, and you should be able to make the calls that follow: precision, quantization, tensor and pipeline and expert parallelism, memory strategy, batching, concurrency, topology, and serving runtime. Then defend them, measure them, and operate them. Doing that well is the baseline. What makes the role interesting is the question sitting underneath it. Serving expertise is scarce, slow to acquire, and currently lives in the heads of a small number of people who have done it enough times to have the instinct. We do not accept that it has to stay that way. A great deal of what those people know is reasoning, and reasoning can be written down, tested, and eventually carried out by something other than a person at two in the morning. How far that can be pushed is genuinely open, and you would be one of the people finding out. This is not a research position and it is not an architecture-only role. You will benchmark, deploy, debug, tune, automate, and operate real inference systems — and then carry them in production.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed