About The Position

Production reliability is a board-level priority, and we are placing it in one hands-on leader. We are hiring a Vice President of Engineering to found and lead our Site Reliability Engineering (SRE) platform organization and own the mission end to end: raise observability maturity and materially reduce production incidents across the engineering estate. You will do it by making agentic engineering the core method of the function — using AI agents to run the platform itself, and putting AI agents in the hands of product engineers so they can meet the company's observability requirements and SRE goals with far less manual effort. This is a builder's role with executive visibility. You will stand up a central platform team that treats reliability and observability as a product, deliver an aggressive multi-phase roadmap, and change how thousands of engineers instrument, operate, and take ownership of what they ship — in a large, complex, tech-debt-heavy environment where the winning strategy is paved roads over mandates. Central to that is a two-part AI mandate: making this function fully agent-native, and delivering AI agents to product engineers so they can fulfill their observability requirements and the company's SRE goals with far less manual effort. You will report directly to the SVP of Platform Engineering and partner with leaders across engineering, product, and the executive team. The mission has a named executive sponsor and committed funding; your job is to turn that mandate into outcomes. The AI mandate: two transformations you will lead: Agentic engineering is not a side initiative in this role — it is central to how the function operates and to the value it delivers to the enterprise. You will be accountable for two connected transformations: 1. Transform this function to be fully agent-native Re-found the platform organization around agentic engineering. Agents — not manual toil — should carry the load of instrumenting legacy code, investigating incidents, generating configuration, assisting on-call, and, over time, executing guarded remediation. You will build the agent control plane, guardrails, and evaluation harnesses that make this safe, and turn the function into the company's proof point for what disciplined, agent-first engineering looks like at scale. 2. Put AI agents in the hands of product engineers for observability and SRE outcomes Deliver AI agents to product engineers as part of the paved road so they can meet the company's observability requirements and SRE goals with far less manual effort — agents that instrument their services, generate SLOs, dashboards, and alerts, investigate incidents, and assist on-call. These agents, and the reliability, evaluation, guardrails, and cost governance behind them, are how product teams hit reliability targets at scale. This role owns the SRE platform: the agents it provides serve observability and reliability, not product-feature development. What you'll own: The SRE platform organization — a central platform team plus embedded/partner SREs, an enablement function, and a cross-cutting reliability champions network. The observability & reliability platform — telemetry pipeline (OpenTelemetry), metrics/logs/traces backends, dashboards and alerting, the SLO and error-budget system, incident management, and RCA/correlation. The paved road — shared instrumentation SDKs, templates, and dashboards/alerts/SLOs-as-code that make golden-signal observability near-automatic. The agentic engineering stack — internal agents for instrumentation, investigation/RCA, config generation, on-call, and guarded remediation — plus the agent control plane, guardrails, and evaluation that keep them safe. The agentic paved road for product engineers — AI agents delivered to product teams to instrument services, generate SLOs/dashboards/alerts, investigate incidents, and assist on-call — so they meet observability requirements and SRE goals with minimal manual effort, backed by guardrails and evaluation. Reliability & AI governance — error-budget policy, blameless incident and postmortem practice, AI safety and human-in-the-loop controls, and the metrics reported to senior leadership. The roadmap, budget, and vendor strategy — an 18-month phased plan, the operating budget, tooling selection, and cost governance for both telemetry and AI workloads. Key responsibilities: Set the strategy and vision. — Own the enterprise strategy for reliability, observability, and agentic engineering. Define what “reliability as a product” and “agent-native engineering” mean here, and keep both tied to business outcomes. Build and lead the organization. — Recruit, structure, and grow a high-performing platform organization spanning SRE, platform, and AI engineering — hiring and developing senior, staff, and principal talent and the managers who lead them. Deliver the roadmap. — Execute the phased plan — mobilize and instrument, build foundations, scale and standardize the paved road with SLO coverage and error-budget policy, then bring proactive and agentic capabilities to production — on aggressive, overlapping timelines. Make the function agent-native. — Drive adoption of internal agents across the platform's own work, with least-privilege access, blast-radius limits, human-in-the-loop controls, and evaluation before any increase in autonomy. Put agents in product engineers' hands. — Deliver AI agents through the paved road that let product engineers meet observability requirements and SRE goals — instrumenting services, standing up SLOs, and resolving incidents — with far less manual effort, backed by guardrails and evaluation. Make reliability measurable. — Stand up SLOs and error budgets, modern incident management, and blameless postmortems; establish error-budget policy in partnership with leadership. Drive adoption across the estate. — Win teams over with paved roads and lighthouse wins, not mandates — through reliability reviews, enablement, office hours, and published before/after results. Own the tooling and its economics. — Select and evolve an industry-leading, OpenTelemetry-native and AI-native tool stack, with cost and cardinality governance built in from the start for both telemetry and inference. Operate as an executive partner. — Manage stakeholders across engineering and product, report progress and reliability/AI metrics to the CTO and executive team, and steward budget and headcount. What success looks like: First 90 days — organization mobilized and sponsor alignment confirmed; reliability baseline established; 2–4 lighthouse services selected; instrumentation underway with the first golden-signal dashboards and SLOs live; the first internal agents piloted. By 12 months — the paved road is self-service and adopted by tier-1 teams; SLOs and error-budget policy are in effect; on-call load and alert noise are measurably down; the function is operating agent-first; and product engineers are using the SRE agents to instrument and meet SLOs with far less manual effort. By 18 months — proactive and guarded agentic capabilities are in production; the program hits its targets — a significant reduction in Sev1/Sev2 incidents, mean-time-to-resolution cut substantially, full SLO coverage on tier-1 services, and a healthier on-call — and the company's SRE goals are being met at scale through agents that product engineers rely on.

Requirements

  • 15+ years in software engineering, including 7+ years in senior engineering leadership leading SRE, platform, infrastructure, or AI organizations at scale — including managing managers.
  • Proven track record of improving reliability and reducing production incidents across a large, complex, multi-team estate.
  • Demonstrated experience building and operating production AI and/or agentic systems at scale, including LLMs, agent frameworks, retrieval, evaluation, guardrails, and AI/agent observability.
  • Experience establishing AI governance and safety for production agents (least-privilege access, human-in-the-loop controls, blast-radius limits, evaluation-gated autonomy).
  • Deep expertise in observability (OpenTelemetry; metrics, logs, traces), SLOs and error budgets, incident management, and cloud-native infrastructure (Kubernetes, infrastructure-as-code).
  • Experience running an internal platform as a product (paved roads, developer experience, adoption measured by usage).
  • Demonstrated ability to drive org-wide change through influence in a large, matrixed organization.
  • Excellent executive communication skills, able to translate strategy into business terms and present to C-level stakeholders.
  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience.

Nice To Haves

  • Track record of building internal AI/agent developer tooling adopted at scale.
  • Experience in financial services or another regulated, high-availability, high-compliance environment.
  • Familiarity with a modern reliability and AI stack (e.g., OpenTelemetry, Prometheus/Grafana, PagerDuty, SLO tooling, LLM/agent observability platforms).
  • Track record establishing error-budget policy and a genuinely blameless postmortem culture.
  • Advanced degree in a relevant field.

Responsibilities

  • Set the strategy and vision for reliability, observability, and agentic engineering.
  • Define "reliability as a product" and "agent-native engineering" tied to business outcomes.
  • Recruit, structure, and grow a high-performing platform organization.
  • Execute the phased roadmap for reliability and agentic capabilities.
  • Drive adoption of internal agents for platform work with safety controls.
  • Deliver AI agents to product engineers for observability and SRE goals.
  • Establish SLOs, error budgets, incident management, and blameless postmortems.
  • Drive adoption across the engineering estate through influence and lighthouse wins.
  • Own and evolve the reliability and AI tool stack, including cost governance.
  • Manage stakeholders, report metrics to leadership, and steward budget and headcount.

Benefits

  • Online Privacy Notice
  • EEO Statement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service