Production reliability is a board-level priority, and we are placing it in one hands-on leader. We are hiring a Vice President of Engineering to found and lead our Site Reliability Engineering (SRE) platform organization and own the mission end to end: raise observability maturity and materially reduce production incidents across the engineering estate. You will do it by making agentic engineering the core method of the function — using AI agents to run the platform itself, and putting AI agents in the hands of product engineers so they can meet the company's observability requirements and SRE goals with far less manual effort. This is a builder's role with executive visibility. You will stand up a central platform team that treats reliability and observability as a product, deliver an aggressive multi-phase roadmap, and change how thousands of engineers instrument, operate, and take ownership of what they ship — in a large, complex, tech-debt-heavy environment where the winning strategy is paved roads over mandates. Central to that is a two-part AI mandate: making this function fully agent-native, and delivering AI agents to product engineers so they can fulfill their observability requirements and the company's SRE goals with far less manual effort. You will report directly to the SVP of Platform Engineering and partner with leaders across engineering, product, and the executive team. The mission has a named executive sponsor and committed funding; your job is to turn that mandate into outcomes. The AI mandate: two transformations you will lead: Agentic engineering is not a side initiative in this role — it is central to how the function operates and to the value it delivers to the enterprise. You will be accountable for two connected transformations: 1. Transform this function to be fully agent-native Re-found the platform organization around agentic engineering. Agents — not manual toil — should carry the load of instrumenting legacy code, investigating incidents, generating configuration, assisting on-call, and, over time, executing guarded remediation. You will build the agent control plane, guardrails, and evaluation harnesses that make this safe, and turn the function into the company's proof point for what disciplined, agent-first engineering looks like at scale. 2. Put AI agents in the hands of product engineers for observability and SRE outcomes Deliver AI agents to product engineers as part of the paved road so they can meet the company's observability requirements and SRE goals with far less manual effort — agents that instrument their services, generate SLOs, dashboards, and alerts, investigate incidents, and assist on-call. These agents, and the reliability, evaluation, guardrails, and cost governance behind them, are how product teams hit reliability targets at scale. This role owns the SRE platform: the agents it provides serve observability and reliability, not product-feature development.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Executive