Staff Operational Support Engineer L2

Avispa TechnologyAtlanta, GA
Onsite

About The Position

A leading video, audio, and voice technologies company is seeking a Staff Operational Support Engineer L2 to provide operational support for 24/7 live video streaming, advertising, player, and real-time delivery platforms. This role involves troubleshooting complex production incidents, performing configuration changes in production environments, and improving operational efficiency through automation and AI.

Requirements

  • 5+ years of relevant experience in operational, support, or similar customer-facing roles.
  • Experience supporting production video streaming platforms, OTT services, and live systems.
  • Troubleshooting skills across distributed systems, including APIs, microservices, and cloud infrastructure.
  • Familiarity with HLS, DASH, CMAF, WebRTC, DRM, and CDN architectures.
  • Experience using monitoring, alerting, and logging tools such as Grafana, Kibana/ELK, Prometheus, and Loki to diagnose live incidents.
  • Ability to correlate backend streaming metrics, player telemetry, and CDN signals to diagnose live customer issues end-to-end.
  • Comfort performing controlled changes in production environments.
  • Working knowledge of incident management and on-call operations.

Responsibilities

  • Own escalated customer issues from Level 1 Support through resolution, troubleshooting complex production incidents affecting live streams, VOD playback, ad insertion, DRM, and real-time WebRTC services.
  • Operate directly in production environments to perform configuration changes, CDN adjustments, mitigations, and emergency changes when required, while providing clear and timely customer-facing communication and leading or contributing to live incident bridges with customers, internal teams, and partners.
  • Work with Infrastructure as Code as the primary mechanism for safe, auditable, and repeatable production changes, using Terraform, Helm, Kubernetes manifests, GitOps workflows, CI/CD and deployment pipelines.
  • Validate and execute infrastructure and configuration changes through codified workflows and collaborate with Engineering and DevOps to improve deployment reliability and operational safety.
  • Improve operational efficiency and incident response, including AI-assisted incident triage and classification, automated runbook execution, AI-based incident pattern detection, intelligent alert correlation and noise reduction, automated or improved incident communications, accelerated troubleshooting workflows, and identification of recurring or systemic issues.
  • Drive adoption of automation-first and AI-augmented operational practices.
  • Support pre-event operational readiness for critical customer events through runbook checks, monitoring coverage validation, risk identification and mitigation planning, and rehearsed incident-response strategies.
  • Respond to critical alerts within defined SLAs for stream health, player errors, and delivery infrastructure.
  • Perform and contribute to root cause analyses, document findings and corrective and preventive actions, identify recurring issues and partner with Engineering and Product teams to eliminate them.
  • Improve runbooks, operational playbooks, and knowledge bases across player, advertising, live-streaming, and real-time products.
  • Support production deployments and defect resolution.
  • Provide feedback on observability, tooling gaps, and operational risks.
  • Serve as the operational voice during post-incident reviews.

Benefits

  • Group Medical
  • Dental
  • Vision
  • Life
  • Retirement Savings Program
  • PSL
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service