Software Engineer II

The Walt Disney Company•New York, NY
•Onsite

About The Position

Disney Entertainment and ESPN Product & Technology is a global organization focused on building and advancing the technological backbone for Disney’s media business. The team combines technology with creativity to create world-class products, enhance storytelling, and drive innovation and scalability. This role is within the Observability & Insights group, which ensures Disney Streaming’s distributed systems are reliable, performant, and transparent by building telemetry, dashboards, alerting, insights pipelines, and developer experience tooling. As a Software Engineer II, you will design and build intelligent, AI-driven systems to enhance the reliability and compliance of Disney's large-scale streaming infrastructure. You will develop agentic systems, machine learning models, and real-time automation to transform telemetry, logs, and infrastructure signals into automated detection, root cause analysis, and proactive policy enforcement. You will contribute to autonomous agents capable of reasoning over complex system behavior, identifying issues in real time, and driving faster detection and resolution across Disney+, Hulu, and ESPN. Additionally, you will contribute to systems that provide infrastructure visibility, enforce compliance, correlate services to incidents, and quantify the impact of AI-driven observability tools. You will collaborate with engineering, product, and platform teams to embed intelligence into operational workflows and improve system resilience and compliance at scale. In this fast-paced, AI-native SRE engineering environment, you will deliver high-quality features end-to-end, contribute to system design and architecture reviews, and own components of production systems.

Requirements

  • 3+ years of applicable experience in backend development, including building AI-powered or data driven applications and scalable APIs (e.g. FastAPI, Flask)
  • Knowledge of AWS services and cloud architecture to inform what infrastructure characteristics should be measured and evaluated for reliability and compliance
  • Practical experience in AI/ML engineering, with knowledge in at least one of the following areas: Agentic Workflows: Orchestrating foundation models (GPT-4, Claude) using frameworks like LangChain or LangGraph. Traditional ML: Developing, training, or fine-tuning models using frameworks like PyTorch or TensorFlow.
  • Proficiency with AI-assisted development tools (e.g., Cursor, Claude Code) to accelerate engineering velocity.
  • Experience with modern development practices, including version control (GitHub), containerization (Docker), and cloud-native deployments (AWS/EKS).
  • Strong understanding of API design, microservices architecture, and standard SDLC workflows.
  • Strong analytical and technical skills to troubleshoot issues, perform rapid iteration and quickly come-up with the possible solutions
  • Strong collaboration and communication skills, with the ability to work cross-functionally and clearly explain complex technical concepts

Nice To Haves

  • Experience with observability platforms (e.g., Datadog, Grafana, Conviva) and handling high-volume telemetry data.
  • Familiarity with large-scale data platforms and distributed data processing tools (e.g., PySpark, Pandas, Databricks, Snowflake).
  • Knowledge of prompt design, model evaluation, and fine-tuning foundation models (GPT-4, Claude).
  • Experience implementing production-grade systems at scale within a fast-paced, distributed environment.

Responsibilities

  • Design and operate intelligent, production-grade systems that leverage real-time signals and AI-driven detection to improve the health of streaming platforms, critical services, and customer experience
  • Build and scale AI-driven capabilities including agentic AI systems powered by modern foundation models (e.g. Claude Opus/Sonnet, GPT-4) enabling automated reasoning and decisioning, as well as predictive modeling and anomaly detection for real-time system health and reliability
  • Develop end-to-end data and decisioning pipelines that transform telemetry, logs, and user signals into actionable insights, automated detection, and root cause analysis
  • Create and deploy scalable APIs and services that deliver predictive signals, explainability, and insights to engineering teams, operational tools, and product stakeholders
  • Contribute to infrastructure visibility and compliance enforcement systems that aggregate resource data across Disney products, enable ownership-scoped querying, and prevent non-compliant deployments
  • Partner cross-functionally to embed intelligence into workflows (incident response, release validation, customer insights), improving speed and reducing operational overhead
  • Drive innovation in observability, reliability, and developer productivity through applied AI and new approaches

Benefits

  • A bonus and/or long-term incentive units may be provided as part of the compensation package, in addition to the full range of medical, financial, and/or other benefits, dependent on the level and position offered.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service