Principal Software Engineer - Observability

The Walt Disney CompanyGlendale, CA
$184,300 - $258,900Hybrid

About The Position

Disney Entertainment and ESPN Product & Technology is seeking a Principal Software Engineer specializing in Observability. This role is part of a global organization focused on building and advancing the technological backbone for Disney’s media business. The team combines technology with creativity to develop world-class products, enhance storytelling, and drive innovation and scalability. The Product Engineering team is responsible for the engineering of Disney Entertainment & ESPN digital and streaming products and platforms, including product engineering, media engineering, quality assurance, and the engineering behind personalization, commerce, lifecycle, and identity. The Observability & Insights group specifically ensures that Disney Streaming’s distributed systems are reliable, performant, and transparent by building telemetry, dashboards, alerting, insights pipelines, and developer experience tooling. This Principal Software Engineer will be a hands-on builder, working across unfamiliar codebases and teams to accelerate software engineering velocity with AI. They will design and build intelligent systems to improve the reliability and performance of Disney’s large-scale streaming ecosystem, utilizing AI for automated detection, root cause analysis, and proactive insights across Disney+, Hulu, and ESPN. The role requires operating with independence, a strong innovation bias, and the ability to tie work to clear business value and security approvals. The engineer will leverage frontier AI models as a force multiplier, optimize costs, and act as a technical thought leader to raise the engineering bar.

Requirements

  • Bachelor’s degree in computer science, Engineering, or equivalent experience.
  • 10+ years of software engineering experience, including building AI-powered or data-driven applications and scalable APIs (e.g. FastAPI, Flask) and deploying production systems at scale.
  • Proven ability to ramp quickly on unfamiliar codebases and rebuild or modernize systems with AI at significantly higher productivity than a typical engineer, while maintaining quality. Tangible examples where your application of AI has led to 2x to 4x times productivity gains are a strong plus.
  • Demonstrated expertise in model prompting, context engineering and harnessing, and the design of orchestration patterns and fully autonomous agentic workflows.
  • Strong AI/ML engineering experience, including orchestrating foundation models (e.g. Claude, OpenAI, Qwen) using frameworks like LangChain or LangGraph.
  • Track record building and deploying high-quality systems across both front end and back end on hyperscaler cloud platforms (AWS, Azure, or GCP), and integrating frontier model APIs (Anthropic, OpenAI).
  • Fluency with AI-assisted development tools (e.g., Cursor, Claude Code) to accelerate engineering velocity, and demonstrated experience optimizing AI usage for cost and performance through prompt optimization, caching, and smart model routing.
  • Experience with modern development practices, including version control (GitHub), containerization (Docker), cloud-native deployments (AWS/EKS), and mature CI/CD pipelines.
  • Strong understanding of API design, microservices architecture, and standard SDLC workflows.
  • Track record of working with extreme independence and innovation, scoping and delivering high-impact work that maps directly to business outcomes, and the ability to set technical direction and mentor other engineers.
  • Strong analytical and technical skills to troubleshoot issues, iterate rapidly, and quickly arrive at viable solutions.
  • Strong collaboration and communication skills, with the ability to work cross-functionally and clearly explain complex technical concepts to technical and non-technical stakeholders.

Nice To Haves

  • Experience building SaaS solutions in the observability space and handling high-volume telemetry data (e.g., Datadog, Grafana, Conviva). Familiarity with OpenTelemetry (OTel) is a bonus.
  • Experience integrating with enterprise AI services such as AWS Bedrock, including model invocation, routing, and governance integration.
  • Experience deploying and serving open-source models hosted locally or in private infrastructure, including inference optimization and cost/performance tuning.
  • Familiarity with large-scale data platforms and distributed data processing tools (e.g., PySpark, Pandas, Databricks, Snowflake).
  • Knowledge of prompt design, model evaluation, and fine-tuning foundation models (e.g., Claude, GPT).
  • Experience implementing production-grade systems at scale within a fast-paced, distributed environment.

Responsibilities

  • Design and operate intelligent, production-grade systems that use real-time signals and AI-driven detection to improve the health of streaming platforms, critical services, and customer experience.
  • Build and scale fully autonomous agentic systems powered by modern frontier models (e.g. Anthropic, OpenAI, Google, Meta) that reason over complex system behavior, decide and act with minimal human intervention, and drive faster detection and resolution.
  • Design intelligent harnesses, memory, and multi-agent orchestration patterns – including prompting and context strategies, task decomposition, retrieval, tool routing, and error recovery – that maximize model performance and accuracy and reduce hallucinations across AI-generated context and actions.
  • Develop end-to-end systems across front end and back end, including data and decisioning pipelines, and scalable APIs that deliver predictive signals, explainability, and insights to engineering teams and product stakeholders.
  • Keep AI usage cost-effective through prompt and context optimization, caching, batching, and smart model routing across frontier APIs and self-hosted open-source models.
  • Apply software engineering best practices end to end: clean, well-tested code, thorough code reviews, and mature CI/CD, owning components of production systems.
  • Drop into brand-new or unfamiliar codebases, rapidly build a working mental model, and modernize them with AI, partnering cross-functionally to embed intelligence into workflows such as incident response, release validation, and customer insights and to drive innovation in observability, reliability, and developer productivity.

Benefits

  • A bonus and/or long-term incentive units may be provided as part of the compensation package, in addition to the full range of medical, financial, and/or other benefits, dependent on the level and position offered.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service