Staff Site Reliability Engineer, Playout

NBCUniversalStamford, CT
$145,000 - $175,000Hybrid

About The Position

NBCUniversal Operations & Technology is looking for a Staff SRE, Playout Engineering to provide technical leadership to a team of Site Reliability Engineers. This team drives reliability, observability, and operational excellence for cloud-based master control playout systems supporting all NBCUniversal live linear channels – NBC, Telemundo, Peacock Virtual Channels, etc. In this position, you will shape the reliability strategy for live linear playout—defining service levels, improving resiliency, and strengthening monitoring and incident response—so the platform can meet evolving business needs with predictable performance and availability. This role requires the ability to operate in a fast-paced environment. For systems in production, you will lead an on-call team and drive L1 and L2 troubleshooting, incident management and continuous improvement to maintain reliable distribution.

Requirements

  • Bachelor's degree in computer science or related degree / experience
  • Eight years’ hands-on-keyboard Engineering experience working with broadcast automation playout environments e.g. Snell, Harris, Imagine, Amagi
  • Requires on-call 24/7 availability for escalations
  • Hands-on-keyboard experience administrating Linux environments
  • Experience with monitoring/logging tools e.g. Splunk and Grafana
  • Experience with streaming protocols and codecs (e.g. TS, HEVC, H.264, HLS, CMAF, SCTE-35, SCTE-224, ESAM, SRT/RIST)
  • Experience with IP networking and interfacing with cloud-based networks
  • Experience with containerization (Docker & Kubernetes)
  • Excellent communicator and able to clearly articulate complex issues and technologies
  • Expert with broadcast playout systems (master control) technologies
  • Expert with public cloud environments using AWS services
  • Comfortable working in a fast-paced agile environment. Requirements change quickly and our team needs to constantly adapt to meet objectives
  • An automate-first and automate everything attitude

Nice To Haves

  • Experience with cloud native playout vendor solutions (Amagi, Evertz, GrassValley, Harmonic, Imagine, CoralBay, Veset, etc.)
  • Experience building strong operational readiness practices (runbooks, alert tuning, on-call health, incident reviews)
  • Ability to create user interface designs based on client workflows

Responsibilities

  • Team lead for SRE engineers on playout Engineering team
  • Define and manage reliability targets (SLIs/SLOs) and operational readiness criteria for playout services
  • Drive incident response: establish on-call practices, lead major incident management, and ensure post-incident reviews result in measurable improvements
  • Partner with engineering, product, and operations teams to improve reliability through capacity planning, performance tuning, and resilience testing
  • Provide high-level conceptual drawings and operational runbooks to support architecture reviews, support readiness, and project planning
  • Leadership in driving automation to reduce toil and improve reliability and mean time to recovery (MTTR)
  • L1 & L2 support to maintain playout infrastructure/services for NBCUniversal including providing after hours on-call support select week
  • Leadership in creating monitoring dashboards (Grafana) and proper alerts(teams/slack/ServiceNow)

Benefits

  • Equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service