AI Infrastructure Engineer

Percepta•New York, NY
•$180,000 - $400,000•Onsite

About The Position

We're hiring an AI Infrastructure Engineer to own the infrastructure, deployment, and operational reliability that powers Percepta's AI systems, including the autonomous agents at the core of what we ship. Part of the work is hardening what exists: tightening our Terraform footprint, strengthening deployment pipelines, bringing more rigor to how we manage infrastructure across regions and providers. Part of it is building what's missing. And part of it is genuinely new territory, figuring out what SRE means when the systems you're operating make autonomous decisions. The infrastructure patterns for the agentic systems of the future don't exist yet. You'll help define them.

Requirements

  • 5+ years operating production infrastructure - SRE, platform, or production engineering.
  • You've operated software in environments you don't control: BYOC, single-tenant, on-prem, or air-gapped. You know what changes when you can't just open the console.
  • Cloud-agnostic by instinct and comfortable in all three of AWS, Azure, and GCP - networking, IAM, managed Kubernetes (EKS, AKS, GKE), and the operational differences that actually bite. Deep in one is table stakes; we need someone who treats the other two as first-class, not as ports.
  • Kubernetes in production on managed clusters across multiple providers, and the judgment to know when a cloud-native service is the right call versus a portable one.
  • Deep Terraform - module design, state, testing, drift, and the judgment to know which blast radius is acceptable.
  • You've carried a pager and then built the thing that made it quieter: incident command, postmortems, SLOs that changed someone's behavior.
  • You've worked under a regulated posture — HIPAA, SOC 2, PCI, FedRAMP — where controls had to live in the infrastructure, not a spreadsheet.
  • Comfortable customer-facing. You will sit in a customer's architecture review and ask their network team for a firewall exception.
  • Python, Go, or Bash, and the instinct to automate a process the second time you do it.
  • Genuine curiosity about what you're operating. Agents, not just pods.

Nice To Haves

  • You've operated infrastructure through a vendor control plane (Ryvn, Nuon, Replicated, or similar) and have opinions about the tradeoff.
  • GitOps and progressive delivery across a fleet of tenants running different versions.
  • Multi-region experience, and a view on where provider-native beats portable in a BYOC product.
  • Real depth in the Grafana stack — Mimir, Loki, and Tempo at multi-tenant scale.
  • GPU and inference operations: Ray, SkyPilot, vLLM or SGLang, LLM gateways and cost attribution.
  • MLOps or research-to-production handoff experience. Not required — you'll be near that boundary either way.
  • You've thought about what observability means for non-deterministic systems, and what a blast-radius control looks like for something that acts on its own.

Responsibilities

  • Build the operational floor: SLOs, alert routing that actually reaches a human, an on-call rotation, incident response, and postmortems that produce changes.
  • Own and evolve the IaC - Terraform across AWS, Azure, and GCP, plus the blueprints and per-environment installations that drive it. Bring it tests, policy-as-code, and drift discipline.
  • Make standing up a new customer environment boring: automate the firewall exceptions, DNS delegations, deploy identities, tag policies, and cert chains that are a runbook today.
  • Keep the three clouds at parity. A service that exists on Azure but not GCP is a half-shipped service, and we don't get to pick which cloud a customer already bought.
  • Make HIPAA and SOC 2 properties of the infrastructure rather than projects: controls in code, evidence generated by the pipeline.
  • Own the observability stack we already run — Grafana, Loki, Mimir, Tempo, Alloy, Langfuse, LiteLLM — and make it answer operational questions about agent behavior: which agent decided what, on whose data, at what cost.
  • Give the engineers who ship into these tenants a paved road, so deploying doesn't require knowing our blueprint template engine.

Benefits

  • competitive compensation
  • equity
  • 401(k) matching
  • comprehensive health and family benefits
  • lunch every day
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service