Engineer, Inference

Sierra•San Francisco, CA
•$230,000 - $390,000•Onsite

About The Position

Sierra is the leading platform for customer-facing AI agents, working with many of the world's biggest brands to transform how they serve customers and grow their businesses. We are primarily an in-person company based in San Francisco, with growing offices across North America, Europe, and Asia. We are guided by a set of values that are at the core of our actions and define our culture: Trust, Customer Obsession, Craftsmanship, Intensity, and a commitment to balancing Family along the way. Sierra’s AI agents depend on foundation models to reason and act in real time. The Inference team builds the systems that make those models fast, reliable, and efficient at scale. As a Software Engineer on Inference, you’ll help define Sierra’s inference architecture across both self-hosted models and third-party inference providers. You’ll work on the systems responsible for serving and routing inference, managing capacity and quota, and optimizing for latency, reliability, and cost. This is a systems-first role at the intersection of distributed infrastructure and AI. We're looking for engineers who love complex systems problems and are excited to apply that expertise to one of the fastest-moving areas of AI infrastructure.

Requirements

  • Deep systems thinking and strong distributed systems fundamentals.
  • Experience designing, building, and operating large-scale production systems.
  • Strong judgment around tradeoffs involving latency, reliability, capacity, and cost.
  • Experience taking ownership of complex infrastructure from architecture through production operation.
  • Excitement about applying systems expertise to AI infrastructure and learning quickly as the underlying technology evolves.

Nice To Haves

  • Experience with ML infrastructure, MLOps, or production inference systems.
  • Experience serving LLMs or other large models at scale.
  • Experience operating self-hosted inference and GPU infrastructure.
  • Familiarity with inference frameworks such as vLLM or SGLang.
  • Experience with post-training infrastructure or inference-performance optimization.

Responsibilities

  • Partner with frontier labs and providers.
  • Shape Sierra’s inference architecture.
  • Build for low latency and high reliability.
  • Build and operate self-hosted inference.
  • Optimize inference performance.
  • Build across a hybrid inference stack.
  • Push the serving stack forward.
  • Support the broader model lifecycle.

Benefits

  • Flexible (unlimited) paid time off
  • Medical, dental, and vision benefits for you and your family
  • Life insurance and disability benefits
  • Retirement plan dependent on country of employment
  • Parental leave
  • Fertility and family building benefits through Carrot
  • Lunch, as well as delicious snacks and coffee to keep you energized
  • Discretionary benefit stipend giving people the ability to spend where it matters most
  • Free alphorn lessons
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service