Software Engineer II

The Walt Disney CompanyGlendale, CA
$117,500 - $165,000Remote

About The Position

Technology is at the heart of Disney’s past, present, and future. Disney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – all working to build and advance the technological backbone for Disney’s media business globally. The team marries technology with creativity to build world-class products, enhance storytelling, and drive velocity, innovation, and scalability for our businesses. We are Storytellers and Innovators. Creators and Builders. Entertainers and Engineers. We work with every part of The Walt Disney Company’s media portfolio to advance the technological foundation and consumer media touch points serving millions of people around the world. Product Engineering is a unified team responsible for the engineering of Disney Entertainment & ESPN digital and streaming products and platforms. This includes product engineering, media engineering, quality assurance, engineering behind personalization, commerce, lifecycle, and identity. The Observability & Insights group ensures that Disney Streaming’s distributed systems are reliable, performant, and transparent. We build ML-powered detection systems, telemetry pipelines, intelligent alerting, and developer experience tooling that enable engineers across the organization to understand system health and take action quickly.

Requirements

  • 3+ years of professional software engineering experience building, scaling, and maintaining ML-powered backends, data-driven microservices, and production RESTful APIs using FastAPI or Flask
  • Strong hands-on experience in end-to-end ML engineering using PyTorch or TensorFlow spanning model architecture selection, feature engineering, training, and evaluation (e.g., autoencoders, sequential/time-series models like RNNs/GRUs, anomaly detection, or transformers).
  • Proficiency in Python and at least one ML framework (PyTorch preferred)
  • Experience managing model experiments, lineage, and hyperparameter tracking using tools like MLflow or Weights & Biases.
  • Practical experience processing, transforming, and querying large-scale telemetry, event, or time-series datasets using PySpark, Pandas, or Databricks.
  • Experience with modern development practices including version control (GitHub), containerization (Docker), and cloud-native deployments (AWS/EKS)
  • Strong analytical and troubleshooting skills with the ability to iterate rapidly
  • Strong collaboration and communication skills, with the ability to work cross-functionally

Nice To Haves

  • Experience with foundation model integration, prompt engineering/evaluation, RAG architectures, or orchestration frameworks like LangChain or LangGraph
  • Familiarity with observability platforms (e.g., Datadog, Grafana, Conviva) and high-volume telemetry data
  • Experience with large-scale data platforms (Databricks, Spark, Snowflake)

Responsibilities

  • Contribute to production ML models for anomaly detection, including autoencoders, statistical threshold models, and ensemble detection systems that monitor thousands of microservices
  • Develop and improve ML training pipelines using PyTorch on GPU clusters — including feature engineering, model training, threshold calibration, and deployment through MLflow
  • Engineer features from time-series telemetry (error ratios, latency, infrastructure metrics) — implementing windowing, normalization, and data quality safeguards for model consumption
  • Support real-time ML inference systems that run prediction cycles in production — including model serving, detection logic, and alert generation
  • Build AI-driven capabilities using foundation models (Claude, GPT-4) for automated investigation and reasoning over system health signals
  • Create and deploy scalable APIs and services (FastAPI) that deliver ML predictions, health status, and insights to engineering teams and operational tooling
  • Partner cross-functionally to embed ML intelligence into workflows such as incident response and release validation
  • Contribute to model improvement through evaluation, retraining, and threshold tuning

Benefits

  • A bonus and/or long-term incentive units may be provided as part of the compensation package, in addition to the full range of medical, financial, and/or other benefits, dependent on the level and position offered.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service