Staff Infrastructure Engineer

IvoSan Francisco, CA
Onsite

About The Position

Infrastructure Engineers build the foundation for Ivo’s entire platform. Customers are cagey about their contracts, so each customer gets their isolated environment with containers, database, VPC, etc. Things break. Regions go down. Cloud and LLM providers have “incidents.” Customers still expect us to hit our SLAs. This role involves running Kubernetes, building internal tooling for infra management, designing workload isolation strategies, implementing security measures, defining SLOs, managing CI/CD pipelines, maintaining Docker environments, debugging across the stack, and automating repetitive tasks. The ideal candidate will own uptime and lead incident response. This is not a "keep the lights on" role; it involves building the system that keeps the company running and pushing the frontiers of LLM infrastructure.

Requirements

  • Experience living in Kubernetes, understanding cluster architecture, scheduling, networking, and storage primitives.
  • Proficiency in Infrastructure as Code (IaC) and CI/CD tools like Pulumi or Terraform, Docker, and GitHub Actions, preferably in multi-cluster, multi-region setups.
  • Ability to think in terms of failure modes and design systems for resilience.
  • Capability to translate abstract compliance and legal requirements into concrete infrastructure implementations.
  • DevOps/infra expertise with enough full-stack range to debug across backend and frontend.
  • Systematic and relentless debugging skills.
  • Comfort with ambiguity and tracing issues across distributed systems.
  • 8+ years of experience.

Nice To Haves

  • Experience working in a startup environment.
  • Excitement about the adventure of building a company.
  • Excitement about LLMs.

Responsibilities

  • Run Kubernetes, owning multi-cluster, multi-region deployments across AWS/GCP/Azure, with failover and disaster recovery.
  • Build internal tooling for spinning up and managing clusters, ensuring consistent dev-to-prod environments.
  • Design strategies for isolating workloads (ML vs. API traffic) balancing cost, performance, and reliability.
  • Implement security measures such as RBAC, workload identity, secrets management, data residency, and audit trails.
  • Define and enforce SLOs that are real, enforceable, and minimize unnecessary alerts.
  • Manage CI/CD pipelines using GitHub Actions and infrastructure as code with Pulumi.
  • Maintain consistent Docker environments across development, staging, and production.
  • Debug issues end-to-end across APIs, workers, jobs, and frontend builds.
  • Automate repetitive tasks to improve build times, environment cleanliness, and reduce manual toil.
  • Build observability to proactively identify and resolve issues before customers report them.
  • Lead incident response and write informative postmortems.

Benefits

  • Competitive Compensation
  • Equity
  • Relocation and Visa Support
  • Comprehensive medical, dental, and vision plans
  • HSA and FSA accounts
  • Life insurance coverage
  • 401(k) Program
  • Commuter Benefits
  • Unlimited PTO
  • Catered lunch five days a week
  • Premium snacks and coffee
  • In-building gym
  • Dog-friendly environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service