Staff Site Reliability Engineer

IvoSan Francisco, CA
Onsite

About The Position

Infrastructure Engineers build the foundation for Ivo’s entire platform. Customers are cagey about their contracts, so each customer gets their isolated environment with containers, database, VPC, etc. Things break. Regions go down. Cloud and LLM providers have “incidents.” Customers still expect us to hit our SLAs. We’re looking for an Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own uptime, reliability, and performance end-to-end Define and enforce SLI, SLO and SLA targets (and make sure we don’t get paged in the wee hours) Design failover + disaster recovery that actually works in real scenarios Turn data residency requirements into real systems (geo-fencing, regional isolation, etc.) Implement security controls that pass audits and don’t slow the product to a crawl Build observability that answers: what, why and how often it broke ? Lead incident response + write postmortems that make people actually read. We need someone who: Minimum 7 years of experience Thinks in failure modes Designs systems that keep working anyway Can translate “this clause in a contract” into actual infrastructure constraints This isn’t a “keep the lights on” role. You’ll be building the system that keeps the company running. In addition to helping us run a solid, high-performance distributed system, we’d love someone who’s as excited about LLMs as we are. You’d be deeply embedded into the engineering team and highly encouraged to push the frontiers.

Requirements

  • Minimum 7 years of experience
  • Thinks in failure modes
  • Designs systems that keep working anyway
  • Can translate “this clause in a contract” into actual infrastructure constraints

Nice To Haves

  • Experience working in a startup environment is preferred but not required.
  • Are excited about the adventure of building a company!
  • As excited about LLMs as we are.

Responsibilities

  • Own uptime, reliability, and performance end-to-end
  • Define and enforce SLI, SLO and SLA targets (and make sure we don’t get paged in the wee hours)
  • Design failover + disaster recovery that actually works in real scenarios
  • Turn data residency requirements into real systems (geo-fencing, regional isolation, etc.)
  • Implement security controls that pass audits and don’t slow the product to a crawl
  • Build observability that answers: what, why and how often it broke ?
  • Lead incident response + write postmortems that make people actually read

Benefits

  • Competitive Compensation
  • Equity: Meaningful ownership in a company that's scaling fast
  • Relocation and Visa Support: We also offer relocation assistance for successful applicants moving to SF, as well as support for visa and green card applications where applicable.
  • Health & Wellness: Comprehensive medical, dental, and vision plans to suit the needs of you and your family.
  • Flexible Spending & Insurance: Access to HSA and FSA accounts, plus life insurance coverage.
  • 401(k) Program: Save for the future with our 401(k) program.
  • Commuter Benefits: We help make getting to and from the office easier and more convenient.
  • Unlimited PTO: So you can take the time you need to recharge, stay healthy, and bring your best self to work.
  • Office Perks: Enjoy a vibrant Downtown San Francisco office with catered lunch five days a week, premium snacks and coffee, an in-building gym, and a dog-friendly environment.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service