Software Engineer, Inference

LumaRedwood City, CA

About The Position

You'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs. This is large-scale inference systems work: scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. It fits a strong systems engineer comfortable with model serving and Kubernetes at scale. If you want pure modeling rather than the systems that run models, this is firmly the systems side.

Requirements

  • Strong Python and system-architecture skills.
  • Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.
  • Experience with queues, scheduling, traffic control, and fleet management at scale.
  • Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.
  • Familiarity with Redis and S3-compatible storage.

Nice To Haves

  • Modern networking stacks including RDMA (RoCE, InfiniBand, NVLink).
  • High-performance large-scale ML systems (100+ GPUs).
  • CUDA, and FFmpeg or multimedia processing.

Responsibilities

  • Ship new model architectures by integrating them into the inference engine.
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
  • Automate, test, and maintain inference services for maximum uptime and reliability.
  • Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
  • Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service