Senior Manager, Site Reliability Engineering

Invoca•Los Angeles, CA
•$190,000 - $250,000•Remote

About The Position

Invoca is building an agentic development environment, one where software is designed, written, tested, deployed, and operated with AI as a first-class collaborator. Infrastructure is where that gets real or doesn't: agents move fast, and the platform underneath them is what makes that safe. The fastest-growing consumer of infrastructure operations isn't a person clicking through a portal, it's software. That’s what we’re building for. You help shape where it goes next. Getting there starts with shedding weight: finding where we can leverage partners and managed services, building self-service paths engineers and agents can both use safely, and retiring what we shouldn't be running at all. This is not a caretaking role. You'd inherit a capable team with deep knowledge of our systems, being asked to work differently than it has. You bring outside perspective: you've built and run infrastructure teams elsewhere and you arrive with a specific point of view on what great looks like. This position reports to the Senior Director, Infrastructure & Security.

Requirements

  • 5+ years hands-on in an SRE, DevOps, systems, or infrastructure engineering role, and 3+ years directly managing teams in one of those disciplines.
  • Real depth across the platforms we run on: Cloud infrastructure in AWS/GCP, Kubernetes, Infrastructure as code (Terraform) and policy-as-code for governing it, GitOps and continuous delivery (ArgoCD and Atlantis), Observability (Prometheus, Grafana, ELK), Linux (configured via Chef), MySQL
  • Strong opinions about infrastructure design and automation, held with an open mind, and the judgment to use both what exists today and where the industry is heading to guide decisions.
  • Significant experience as a hiring manager — you know what strong looks like at this level because you've hired it — and a track record of inheriting a team and raising its performance through development, hiring, and clear expectations.
  • You can work alongside highly skilled external specialists, able to hold your own technically, and can hold them to a high bar.
  • A working point of view on AI in infrastructure work: where it genuinely accelerates the work and where it doesn't, what output you can trust and what needs a gate, and how to raise a whole team's fluency rather than leaving it to individuals.
  • Effective in a remote-first, asynchronous culture, and able to report complex technical progress in terms a non-engineering audience can act on.

Responsibilities

  • Decide what we run and what we don't. We keep a system in-house only when it differentiates us. Apply that test system by system and find the right home.
  • Land the move from self-hosted to managed Kubernetes making sure the operational expertise stays with us.
  • Reduce what the team carries operationally so capacity goes toward building and not maintenance.
  • Own the technical direction for the platform your team runs, partnering with Developer Enablement, Infosec, and architecture where decisions reach beyond it.
  • Lead a team of 8 SREs with deep systems knowledge. Set a clear bar, give people the feedback and context to hit it, and invest in the growth of the team.
  • Own and execute the rolling 3-month delivery plan: translate infrastructure strategy into clear priorities, navigate ambiguity, and hold full accountability for shipping high-impact work.
  • Represent SRE where infrastructure decisions get made, and be in architecture and planning conversations before your team is needed.
  • Work alongside a leadership team including a tech lead, architect, and your engineering management cohort to solve problems bigger than any one team.
  • Drive a pipeline, not a bottleneck. Work the team does on another team's behalf becomes a callable API with a policy decision attached — invoked by an agent or a person with the same governance and the same audit trail.
  • Extend policy and scanning controls so agent-driven infrastructure change is safe by default rather than safe by review.
  • Define clean boundaries and contracts between infrastructure and the product teams that build on it.
  • Win on adoption, not mandate. The path with policy and audit attached has to also be the fastest one available for an agent or a person.
  • Use AI and agentic workflows in how the team builds and operates infrastructure: agent-assisted development, automated incident detection and triage, self-healing deployments, drift and policy detection.
  • Establish safe, observable, and auditable ways to bring AI into how we build and run infrastructure, and help the teams around you adopt them with confidence.
  • Build shared practices and tooling for infrastructure that every team can adopt, replacing one-off, team-specific approaches.
  • Treat infrastructure as a product with internal users: define adoption goals and measure whether teams are actually better off because of your work.
  • Translate technical decisions in the platform into terms leadership and product teams care about: velocity, cost, risk, customer trust, and revenue impact.
  • Partner with engineering and product leadership to prioritize your work based on measurable business and customer impact, not technical elegance alone.

Benefits

  • Medical, dental, and vision coverage begins on your first day of employment for U.S.-based teammates.
  • Access to mental wellbeing support and an Employee Assistance Program (EAP).
  • Wellness Subsidy – Reimbursement that can be applied toward gym memberships, fitness classes, and more.
  • Flexible Time Off
  • 20 U.S. paid holidays, including a winter break
  • Up to 12 weeks of 100% paid leave for baby bonding, adoption, and caring for family members.
  • Up to 12 weeks of 100% paid leave for childbirth and medical needs.
  • Access to industry-leading enterprise AI tools, including leading LLMs and agentic AI workflow solutions to help you maximize productivity and work more effectively.
  • Annual reimbursement of up to $1,000 per fiscal year to support learning, certifications, conferences, and other professional development opportunities.
  • 401(k) plan through Fidelity with a company match of up to 4%.
  • Recognition Programs
  • Sabbatical after seven years of service.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service