MLOps Engineer

SPREEAISan Francisco, CA
$145,000 - $180,000

About The Position

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

Requirements

  • Experience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)
  • Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment

Nice To Haves

  • You treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident.
  • You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.

Responsibilities

  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale
  • Partner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions
  • Evaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service