Software Engineer, Infrastructure

DescriptSan Francisco, CA
Hybrid

About The Position

Platform owns the foundation the company runs on: compute and deployment, reliability and on-call, CI/CD and monorepo health, developer environments, the infrastructure that model training and inference run on, and the security boundaries around all of it. AI and agent tooling is the clearest example. You will contribute to how models are trained and served here, as well as how agents work inside our codebase: the environments they run in, the verification that makes their output trustworthy, and the review paths that keep it all legible. There is no industry standard and you’ll help form our opinions rather than inheriting one. Our users are other engineers. Expect to spend time with all the other engineering teams, understanding their needs and building a roadmap. The scope is large, so you'll be choosing what to leave alone as much as build. Architecture decisions here last: this is a small team covering a large surface. You'll have real room to decide things, and you'll stay close to the systems you decide about.

Requirements

  • 8+ years building and operating production distributed systems, or equivalent server-side engineering with a heavy infrastructure focus.
  • Effectively leverage agents to multiply your impact and think critically about how and when to harness AI in your work.
  • Experience running systems where failure was expensive, with opinions about reliability and deployment derived from consequences.
  • Experience carrying a pager, commanding an incident, and rolling back systems.
  • Experience using SLOs and error budgets as operating tools.
  • Experience using a major cloud provider and Kubernetes in production, with infrastructure-as-code as your default.
  • Experience owning an architecture or migration, from planning through launch, with long-lasting consequences.
  • Experience finding important unowned work, scoping it, earning buy-in, and delivering it without a spec.
  • Ability to build a minimal repro, read logs, and write targeted checks to prove a fix works instead of trusting output.

Nice To Haves

  • GPU and ML infrastructure: capacity planning, training or inference pipelines, serving cost and latency.
  • Experience with expensive capacity-constrained systems (if not GPU fleet).
  • Production security engineering: IAM, secrets, supply chain, least privilege.
  • Cloud cost modeling: commitment strategy, reservations, unit economics.
  • CI/CD at monorepo scale.
  • Developer-environment work.
  • Intricacies of Video: Media, video, or GPU-backed workloads.
  • Experience in small teams owning a large surface (e.g., Series B to D company or internal platform team).

Responsibilities

  • Own our platform: GCP, Kubernetes, Temporal, the GPU fleet behind cloud export, and the deploy and rollback machinery everything ships through.
  • Be in the on-call rotation, and make it quieter and more actionable.
  • Own the AI enablement substrate: GPU capacity, training and inference pipelines, and the reliability and cost of the systems serving models in production.
  • Make smart trade-offs regarding cost as an engineering constraint.
  • Own the metering and attribution behind usage numbers as inference grows.
  • Manage security for the systems you run: identity and access, secrets management, least-privilege boundaries, and supply-chain integrity.
  • Make what you build legible through self-explaining infrastructure-as-code, runbooks, and in-repo context.
  • Improve how the team learns and ships by forming hypotheses, instrumenting work, releasing incrementally, and reading results honestly.
  • Strengthen the tooling, standards, tests, observability, and release practices that help the team move quickly without compromising quality.
  • Raise the team’s technical ambition by providing architectural direction, thoughtful reviews, mentoring, and clear human writing.

Benefits

  • Base salary: $220,000 to $292,000
  • Equity
  • Generous healthcare package
  • 401k matching program
  • Catered lunches
  • Flexible vacation time
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service