Software Engineer - MLOps

R37 Lab, R1 RCMNew York, NY
$140,000 - $300,000Onsite

About The Position

You’ll own the production runtime for Phare’s ML stack - deploying, serving, and scaling models across inference endpoints and batch/streaming workflows. You’ll build progressive delivery pipelines with automated rollouts and rollbacks, manage SLOs for latency and availability, and instrument end-to-end observability (metrics, logs, traces, drift, regression). You’ll harden the platform with Terraform, Kubernetes, and CI/CD, ensuring reproducible, auditable ML releases. We are hiring across several seniority levels ranging from Mid-level up to Staff. At a minimum, we expect 5 years of software engineering experience with 2 years of ML Ops experience. This is an in-person role in NYC, requiring at least 3 days in the SoHo office.

Requirements

  • 5 years of software engineering experience.
  • 2 years of ML Ops experience.
  • Background in operating ML systems at scale.
  • Experience deploying and operating models running on GPUs in production - APIs and batch/streaming inference.
  • Strong with Docker/Kubernetes.
  • Experience with IaaC (e.g., Terraform).
  • Experience with CI/CD for services and model artifacts.
  • Experience maintaining environment parity, reproducible releases, and robust model/experiment versioning with data lineage.
  • Experience using progressive delivery with automated rollouts/rollbacks.
  • Experience building end-to-end observability (metrics, logs, traces, and model telemetry for drift/regression) plus actionable alerting, runbooks, and incident response.

Nice To Haves

  • Experience in regulated environments (e.g., healthcare, finance).

Responsibilities

  • Deploying, serving, and scaling models across inference endpoints and batch/streaming workflows.
  • Build progressive delivery pipelines with automated rollouts and rollbacks.
  • Manage SLOs for latency and availability.
  • Instrument end-to-end observability (metrics, logs, traces, drift, regression).
  • Harden the platform with Terraform, Kubernetes, and CI/CD, ensuring reproducible, auditable ML releases.
  • Manage model registries and stage gates.
  • Design scheduled or event-driven retraining when appropriate.
  • Enforce RBAC, secrets management, encryption, and audit logs.

Benefits

  • Top-of-market compensation
  • Flexible PTO
  • Comprehensive health benefits
  • 401(k) matching
  • Annual bonus plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service