Manager Site Reliability Engineering

The Walt Disney CompanyOrlando, FL
Onsite

About The Position

This role is not remote and requires the candidate to be in the local area or willing to relocate. The Manager Site Reliability Engineering supports the Parks Commercial Systems and Consumer Products systems within the Commerce Site Reliability Engineering organization. The manager is responsible for a team that performs incident management and requests for work in both Systems Engineering and Site Reliability engineering functions. The manager will work closely with a diverse team of engineers to deliver observability, availability, and security to application teams. The team also collaborates with other leaders to deliver functions for the overall commerce area, including SDLC and AI functions. The leader will be responsible for managing a team of Cast Members, including regular reviews and OKRs.

Requirements

  • Minimum of 8 years of related work experience.
  • Demonstrated leadership in implementing observability principles across complex systems and environments.
  • Extensive experience with a wide range of continuous integration tools, including Gitlab, AWS CodeBuild, CodeDeploy, CodePipeline, and Azure DevOps.
  • Proficiency in designing and managing highly scalable and resilient infrastructure using configuration management and orchestration tools such as Terraform, Cloud Formation, Ansible, and Chef.
  • Leveraging AI for predictive insights.
  • Outstanding communication and leadership abilities.
  • A visionary who motivates teams to excel and fosters creativity.
  • An advocate for a diverse and inclusive culture that encourages innovation.
  • Bachelor’s degree in Computer Science, Information Systems, Software, Electrical or Electronics Engineering, or comparable field of study, and/or equivalent work experience.

Nice To Haves

  • A deep understanding of containerized and serverless architectures and strategies.
  • Experience with Kubernetes.
  • Experience with AWS and / or GCP.
  • Experience with management of a cross-functional team.

Responsibilities

  • Oversee finances and budgets in MyPPM, ensure accurate billing processes, and contribute to forecasting and accrual processes.
  • Lead the evolution of DevOps practices within the broader team framework, guiding others in leveraging this culture to enhance observability practices.
  • Manage the Site Reliability Engineers to deliver monitoring and observability for development and business users.
  • Manage the design, build, and support of products platforms.
  • Drive teams to consult, design, build, and support development pipelines, automate infrastructure and operations, build telemetry for monitoring, engineer high-reliability, and reinforce best practices to secure company data.
  • Lead all aspects of systems administration skills on Google, Amazon, and Azure clouds as well as on-premise systems.
  • Strategize systems administration in RHEL, Bottle rocket, Kubernetes, and containers.
  • Bring knowledge on systems, network, operational excellence, application stability, security, performance, and capacity management.
  • Engage in estimation and planning across the organization, voicing recommendations, feedback, and solutions from a technical perspective and aligning to overall project goals.

Benefits

  • A bonus and/or long-term incentive units may be provided as part of the compensation package.
  • Full range of medical, financial, and/or other benefits, dependent on the level and position offered.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service