Senior AI Inference Engineer

StackYak
•Remote

About The Position

StackYak is building the infrastructure layer for AI. We are an early-stage, funded company building software that brings compute, GPU infrastructure, networking, and inference together into one product. The opportunity is large, the market is moving quickly, and we are building for real production workloads from the start. This is not internal IT. This is not a slow-moving infrastructure team maintaining someone else's platform. The infrastructure is the product. We are a small, senior team with very little bureaucracy. This is a founding role in its discipline. You will be the first person here whose primary responsibility is the systems layer beneath the product, and the shape it takes will largely be yours to decide. Treat this document as a starting point rather than a boundary. The people who do well here take ground early and are not asked to give it back. We move quickly. We do not have months for someone to learn the fundamentals of their discipline. You should already be very good at what you do, be able to ramp into adjacent areas quickly, and be comfortable operating without perfect requirements or neatly defined boundaries.

Requirements

  • Run large language models in production, under real load, with someone depending on them.
  • Serve a model across multiple GPUs, and across multiple nodes.
  • Size a model against hardware and be correct about whether it would fit.
  • Choose a quantization and precision strategy with real quality and cost consequences.
  • Tune batching and concurrency past the point of easy wins.
  • Measure latency, TTFT, throughput, utilization, and cost, and defend the numbers.
  • Work in NVIDIA and/or AMD inference environments.
  • Write Python that other people run in production (tooling and product, not scripts).
  • Debug Linux and systems problems below the container boundary.
  • Find the root cause of an inference failure that was not in the serving runtime.

Nice To Haves

  • Run inference infrastructure at an AI company, inference provider, GPU cloud, hyperscaler, or serious internal AI platform.
  • Work across multiple GPU generations and vendors, and articulate how that changed decisions.
  • Understand distributed inference and the networking implications of multi-node serving.
  • Benchmark and compare serving runtimes.
  • Explain why a deployment is configured a certain way, rather than repeating vendor recipes.
  • Automate deployment decisions, or build schedulers, placement systems, capacity planners, or similar infrastructure.
  • Build and test things outside the assigned roadmap to understand how they work.

Responsibilities

  • Determine how a model should run given a set of GPUs and production requirements, including precision, quantization, tensor/pipeline/expert parallelism, memory strategy, batching, concurrency, topology, and serving runtime.
  • Defend, measure, and operate the chosen inference configurations.
  • Develop reasoning for inference serving that can be documented, tested, and potentially automated.
  • Benchmark, deploy, debug, tune, automate, and operate real inference systems.
  • Carry inference systems in production.
  • Own how models run, including placement, parallelism, precision, and memory strategy across single-GPU, multi-GPU, and multi-node, and document the reasoning.
  • Own the serving runtimes (e.g., vLLM, SGLang, TensorRT-LLM), including selection and adaptation when necessary.
  • Own what the company can promise regarding concurrency limits, latency and TTFT targets, and throughput under load, providing evidence for these numbers.
  • Establish benchmarking as an institution, including methodology, harness, and reproducibility.
  • Manage the model lifecycle in production, including loading, startup, health, upgrades, capacity, and failure handling.
  • Define the boundaries between multi-tenant and dedicated serving.
  • Handle inference incidents, including those not caused by the serving runtime.
  • Translate inference expertise into software to eliminate single points of failure.
  • Produce sensible, measurable, reproducible, and operable deployments quickly.
  • Help make inference expertise programmatic.
  • Participate in technical decisions.
  • Share responsibility for production and on-call.
  • Move between design, implementation, debugging, and operations.
  • Execute technical decisions after arguing the technical case.

Benefits

  • Competitive compensation
  • Meaningful equity
  • Remote and distributed work environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service