About The Position

At Netflix, our mission is to entertain the world. Together, we are writing the next episode - pushing the boundaries of storytelling, global fandom and making the unimaginable a reality. We are a dream team obsessed with the uncomfortable excitement of discovering what happens when you merge creativity, intuition and cutting-edge technology. Come be a part of what’s next. We are looking for a Senior Manager of Site Reliability Engineering to lead one of the most consequential infrastructure organizations at Netflix. This role owns two intersecting mandates: setting the reliability standards that the entire engineering organization builds to, and leading the SRE team supporting our streaming architecture. Netflix’s infrastructure is undergoing a fundamental shift. Infrastructure is quickly evolving towards a millions-of-agents ecosystem, with AI agents increasingly embedded in how we detect, diagnose, and remediate incidents; how we plan capacity; and how we evolve our reliability posture over time.

Requirements

  • 12+ years in software/infrastructure, with 5+ years in senior SRE leadership.
  • Deep fluency in cloud-native scale (AWS/GCP, Containers, Service Mesh) and modern observability (Metrics, Tracing, Logging).
  • Proven ability to drive technical adoption across complex, decentralized organizations through influence rather than mandate.
  • Ability to navigate the "human-machine" boundary of automation and clearly articulate technical trade-offs to non-technical stakeholders.

Nice To Haves

  • Experience in streaming media, ad-tech, or high-scale gaming backends.
  • Hands-on design of LLM-based autonomous agents in production.
  • Familiarity with Netflix’s OSS ecosystem (Spinnaker, Atlas, Mantis) or Chaos Monkey.
  • Demonstrated experience improving SRE AI fluency.

Responsibilities

  • Build and scale a world-class SRE function, defining the operating model for how SREs partner with product and infrastructure teams.
  • Establish and socialize company-wide standards (SLIs/SLOs, Error Budgets) and publish transparent reliability scorecards to drive engineering accountability.
  • Collaborate with CDN, Playback, and Ads teams to eliminate systemic failures and translate technical reliability data into actionable business risk for executives.

Benefits

  • Health Plans
  • Mental Health support
  • a 401(k) Retirement Plan with employer match
  • Stock Option Program
  • Disability Programs
  • Health Savings and Flexible Spending Accounts
  • Family-forming benefits
  • Life and Serious Injury Benefits
  • paid leave of absence programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service