About The Position

We're looking for a highly technical Staff Platform Engineer to lead the platform and security foundations of Maestro — Félix's internal, identity-aware AI teammate. Maestro already runs at meaningful scale (500+ per-user pods on private GKE) and lives where work happens: Slack, incident rooms, and an emerging agentic intranet, acting across our toolchain (GitHub, Google Workspace, ClickUp, Notion, PagerDuty, New Relic). The interesting problems here are not prompts or models. Once an AI teammate can open a pull request, page an engineer, or query production, the hard questions become identity, credentials, isolation, blast radius, and audit. This is a platform and security role for an agentic system — you'll own the secure, multi-tenant runtime that makes delegated AI work safe at scale. You'll be the technical anchor for Maestro's infrastructure and security within the AI team: architecting the control plane, hardening the runtime, running the fleet, and shaping what the platform needs next as adoption grows. AI is the domain you'll operate in — deep platform and security engineering is the craft we're hiring for.

Requirements

  • Experience: 8+ years in software/infrastructure engineering, with a proven track record owning large-scale, security-critical distributed systems end to end.
  • Platform & Kubernetes mastery (Staff bar): Deep, hands-on Kubernetes in production — operators/CRDs and controller-runtime, Helm, runtime isolation (gVisor or equivalent), multi-tenancy, and fleet operations at scale. Strong cloud-native architecture on GCP (or AWS/Azure), and IaC with Terraform.
  • Security & identity depth (Staff bar): Strong applied security engineering — workload identity (SPIFFE/SPIRE, Workload Identity Federation), service mesh mTLS (Istio), OAuth 2.0 / OIDC, token exchange, JIT/short-lived scoped credentials, secrets/KMS envelope encryption, least-privilege and zero-trust patterns, and threat modeling for adversarial workloads (confused-deputy, prompt injection, data exfiltration).
  • Systems & code: Excellent Go and/or Python, with deep system-architecture judgment. Comfortable owning services, operators, and tooling in production.
  • Observability & LLMOps: Production-grade monitoring, tracing, and audit design (OpenTelemetry), plus SRE fundamentals — SLOs, incident response, cost/performance engineering.
  • Agentic AI (Senior / domain level): Solid, hands-on experience with production LLM/agent systems — tool-calling, multi-step orchestration, human-in-the-loop, model routing, and eval/guardrail thinking. You understand agentic architectures deeply enough to secure and operate them; you do not need to be a model researcher.
  • Ownership & leadership: High autonomy in an early-stage squad — independently diagnose bottlenecks, propose architecture, and ship it. Proven ability to grow engineers through architectural guidance, not just code review, and to align technical decisions with stakeholders across Product, Security, and Leadership.

Responsibilities

  • Own the Maestro platform architecture. Design, build, and operate the multi-tenant control plane and per-user runtime on private GKE — the Kubernetes operators (CRDs/controller-runtime), Helm charts, gVisor-sandboxed pods, and per-user isolation primitives (KSA/GSA, Workload Identity, NetworkPolicy, per-user workspaces) that reconcile one user into a fully wired, isolated environment.
  • Lead the security model end to end. Treat the LLM and its tools as adversarial. Own identity separation (requester / actor / persona), JIT short-lived scoped tokens, an encrypted OAuth refresh-token vault (CMEK/Cloud KMS), zero-credential egress patterns, and a policy layer that decides whose credentials an agent uses — never the prompt.
  • Harden service-to-service trust. Enforce mesh identity with Istio mTLS + SPIFFE, signed request claims (JWS) to prevent confused-deputy issues, and deny-by-exception networking across the fleet (Istio AuthorizationPolicy + Kubernetes NetworkPolicy).
  • Operate the fleet, not the bot. Build fleet health, scale-to-zero, resource packing, and safe operational tooling for 500+ pods, with the SRE-grade availability, latency, and recovery the platform demands.
  • Own IaC and delivery. Drive Terraform for the dedicated GCP projects (VPC, private GKE, GPU/gVisor node pools, Cloud SQL, Memorystore, GSM, Artifact Registry), plus CI/CD and progressive delivery for control-plane and runtime components.
  • Make audit a product feature. Own the OpenTelemetry pipeline (logs/metrics/traces) fanning out to Cloud Logging, BigQuery, and New Relic, capturing gateway, kernel/gVisor syscall, and real-time SecOps events so "who asked, which persona acted, which credentials were used, did the user confirm?" is always answerable.
  • Enforce human-in-the-loop and guardrails. Build the approval flows for irreversible actions (writes, merges, admin ops) and the read-only, model-immutable guardrail mounts (identity, instructions, curated skills).
  • Integrate the agent layer, safely. Partner on the OpenClaw gateway, controlled tool wrappers, and model routing (Vertex AI for stakes, self-hosted Ollama for volume) — ensuring every agent capability is a governed capability, not a raw CLI or API key.
  • Set the technical direction. Define platform and security best practices, mentor senior and mid-level engineers, and map Maestro's next infrastructure needs as it scales.

Benefits

  • Competitive salary
  • Initial stock options grant
  • Annual performance bonus
  • Health, dental, and vision plans
  • 401(k) with employer match
  • Continuous learning opportunities
  • Unlimited PTO
  • Paid parental leave
  • Empowering opportunities for growth in a dynamic entrepreneurial environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service