Member of Technical Staff - Model Serving / API Backend Engineer

Black Forest LabsSan Francisco, CA
$180,000 - $300,000Hybrid

About The Position

This role is critical for bridging the gap between frontier research and production reality in a fast-paced generative AI company. The primary goal is to remove the bottleneck in productionizing research checkpoints, ensuring faster API deployment, optimized inference speed, and robust API performance under load. The engineer will own the development and maintenance of high-performance inference services and APIs, optimize GPU infrastructure for latency and throughput, build scalable serving architectures, and improve system reliability and observability. This position involves close collaboration with researchers to rapidly move from idea to live endpoints and prototype demos. The role requires a strong understanding of backend systems, GPU performance, and production ML serving. The ideal candidate will have experience building and operating systems at scale, understanding the difference between research prototypes and production systems, and making informed tradeoffs regarding performance, reliability, and cost. Comfort in a fast-moving, research-adjacent environment and a strong sense of ownership are essential.

Requirements

  • Built and operated systems at meaningful scale
  • Understand the difference between a research prototype and a production system
  • Comfort navigating ambiguity, making tradeoffs, and improving systems under real-world constraints
  • Strong judgment around performance, reliability, and cost tradeoffs
  • Experience scaling APIs or ML systems under load
  • Comfort working in fast-moving, research-adjacent environments
  • Ownership from system design through debugging and deployment
  • Building and operating ML inference services in production
  • Designing scalable API architectures with async processing
  • Optimizing GPU workloads (batching, quantization, compilation, CUDA)
  • Managing distributed systems and task queues under variable load
  • Implementing monitoring and observability for production ML systems
  • Debugging performance bottlenecks across model, infrastructure, and network layers

Nice To Haves

  • Real-time or low-latency inference systems
  • TensorRT, reduced precision, layer fusion, or model compilation techniques
  • Frontend demo tooling (Streamlit, Gradio, React)
  • CI/CD and automated testing for ML systems
  • Security best practices for API and model serving

Responsibilities

  • Turn research checkpoints into production-ready inference services
  • Design and maintain high-performance APIs serving millions of requests
  • Optimize inference latency and throughput across GPU infrastructure
  • Build scalable serving architectures that handle unpredictable traffic
  • Improve reliability, monitoring, and observability across model-serving systems
  • Prototype and ship demos that showcase new capabilities in days, not weeks
  • Collaborate closely with researchers to move from idea to live endpoint rapidly

Benefits

  • Travel costs covered for in-person meetings
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service