Staff Cloud/ML Ops Engineer

Ivo Inc.San Francisco, CA
$287,000 - $485,000Onsite

About The Position

Infrastructure Engineers build the foundation for Ivo’s entire platform. Customers are cagey about their contracts, so each customer gets their isolated environment with containers, database, VPC, etc. Things break. Regions go down. Cloud and LLM providers have “incidents.” Customers still expect us to hit our SLAs. We’re looking for a Cloud/MLOps Engineer as part of Infrastructure team to own and evolve our Kubernetes platform across AWS/GCP/Azure, design and operate multi-cluster / multi-region architectures with failover and disaster recovery strategies adhering to secure cluster isolation boundaries, build internal tooling for cluster provisioning and lifecycle management, standardizing environments (dev → staging → prod), design strategies to isolate ML vs API workloads while optimizing for cost, performance, and reliability, implement security and compliance controls at the platform layer with RBAC, workload identity, secrets management while preserving data isolation aligned with residency requirements and auditability for enterprise customers, and partner with SRE + ML teams to ensure SLOs are realistic and enforceable and models are deployed in production environments. This isn’t a “keep the lights on” role. You’ll be building the system that keeps the company running. In addition to helping us run a solid, high-performance distributed system, we’d love someone who’s as excited about LLMs as we are. You’d be deeply embedded into the engineering team and highly encouraged to push the frontiers.

Requirements

  • Minimum 7 years of experience
  • Deep, hands-on experience with Kubernetes in production (you’ve debugged it at 2am, not just deployed to it)
  • Strong experience with infrastructure as code (Pulumi, Terraform, etc.)
  • Strong understanding of cluster architecture, scheduling, networking, storage primitives and failure modes in distributed systems
  • Experience managing multi-cluster or multi-region setups with Github CI/CD

Nice To Haves

  • Experience working in a startup environment is preferred but not required.

Responsibilities

  • Own and evolve our Kubernetes platform across AWS/GCP/Azure
  • Design and operate multi-cluster / multi-region architectures with failover and disaster recovery strategies adhering to secure cluster isolation boundaries
  • Build internal tooling for cluster provisioning and lifecycle management, standardizing environments (dev → staging → prod)
  • Design strategies to isolate ML vs API workloads while optimizing for cost, performance, and reliability
  • Implement security and compliance controls at the platform layer with RBAC, workload identity, secrets management while preserving data isolation aligned with residency requirements and auditability for enterprise customers
  • Partner with SRE + ML teams to ensure SLOs are realistic and enforceable and models are deployed in production environments

Benefits

  • Competitive Compensation
  • Equity
  • Relocation and Visa Support
  • Comprehensive medical, dental, and vision plans
  • Access to HSA and FSA accounts
  • Life insurance coverage
  • 401(k) Program
  • Commuter Benefits
  • Unlimited PTO
  • Catered lunch five days a week
  • Premium snacks and coffee
  • In-building gym
  • Dog-friendly environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service