About The Position

Seeking a highly experienced Senior Manager, Site Reliability and Production Delivery Management to lead reliability, availability, production deployment, and operational excellence practices across critical technology platforms. This leader will be responsible for defining and executing the strategy for observability, SRE practices, release governance, change management, and production deployment processes to ensure highly available, secure, scalable, and resilient customer-facing applications and services. The role will partner closely with Application Development, Infrastructure, Platform Engineering, Architecture, Operations, and Product Management teams to drive a culture of reliability, automation, continuous improvement, and operational accountability. The successful candidate will establish enterprise standards for monitoring, incident management, release engineering, deployment automation, and service reliability while enabling faster and safer delivery of business capabilities. This leader will drive operational excellence, resiliency, automation, and governance while enabling the safe and efficient delivery of technology solutions.

Requirements

  • Undergraduate degree or Technical Certificate
  • 10+ years related experience
  • Expert knowledge of the business and bank technology standards (e.g., infrastructure, architecture, processes, applications) from a strategic perspective and managing/ directing teams and projects
  • Sound knowledge of external competition, emerging, industry and/or market trends in relation to own business.
  • Understands strategic direction (including financials) and champions alliances to benefit the Bank and/or department
  • Advocates for operational improvements to enhance the divisions value to the organization

Nice To Haves

  • Graduate degree, preferred
  • Experience in Technology Operations, SRE, Platform Engineering, Infrastructure, DevOps, Release Management, or related disciplines.
  • Experience leading managers and technical teams within large enterprise environments.
  • Deep knowledge of: Site Reliability Engineering (SRE)
  • Observability platforms and monitoring frameworks
  • Incident, Problem, and Change Management
  • Release and Deployment Management
  • ITIL processes
  • Cloud and hybrid infrastructure platforms
  • CI/CD and DevOps practices
  • Experience managing large-scale production environments supporting mission-critical customer-facing applications.
  • Proven ability to lead cross-functional teams and drive reliability transformation, deployment automation, and operational excellence initiatives.
  • Technologies & Tools: Observability & Monitoring: Dynatrace, Splunk, Datadog, Grafana, AppDynamics, Prometheus, ELK Stack
  • Logging & Tracing: OpenTelemetry, Distributed Tracing, Centralized Logging
  • Metrics & Reliability: SLOs, SLIs, Error Budgets, DORA Metrics, Capacity Management
  • Cloud Platforms: Azure, AWS, Google Cloud Platform
  • Container Platforms: Kubernetes, OpenShift, Docker
  • Development Frameworks: JAVA, Angular, Database (Oracle, MS SQL, MongoDB), KAFKA
  • CI/CD & DevOps: GitHub, Azure DevOps, Jenkins, Harness, ArgoCD
  • Automation & Scripting: Python, PowerShell, Shell Scripting, Terraform, Ansible
  • Release & Change Management: ServiceNow, Change Advisory Board (CAB), Deployment Runbooks

Responsibilities

  • Lead and mature platform Observability and SRE practices, including monitoring, alerting, SLOs, SLIs, error budgets, and operational readiness.
  • Establish standards for dashboards, logging, tracing, incident response, runbooks, and on-call support.
  • Drive continuous improvement in reliability, resiliency, automation, and operational efficiency.
  • Lead release planning, deployment governance, and change management processes across multiple delivery teams.
  • Oversee production deployments, deployment readiness reviews, risk management, and rollback strategies.
  • Partner with Engineering, Architecture, Infrastructure, Operations, and Product teams to improve service stability and delivery performance.
  • Provide leadership during major incidents, post-incident reviews, and service restoration activities.
  • Track and report key operational metrics, including availability, MTTR, deployment success rate, and change failure rate.
  • Define and execute the platform roadmap for Observability, SRE, Release Management, and Production Deployment.
  • Build and lead high-performing teams responsible for reliability engineering, monitoring, release coordination, and deployment governance.
  • Establish service reliability objectives, operational standards, and performance metrics aligned with business goals.
  • Champion a reliability-first culture across technology delivery and operations organizations.
  • Provide executive-level reporting on service health, availability, deployment performance, and operational risks.
  • Develop and mature SRE practices, including SLOs, SLIs, error budgets, capacity planning, resiliency engineering, and operational readiness reviews.
  • Drive adoption of automation to reduce operational toil and improve service stability.
  • Lead incident management, post-incident reviews, root cause analysis, and continuous improvement initiatives.
  • Ensure application teams incorporate reliability, recoverability, and non-functional requirements throughout the software development lifecycle.
  • Establish standards for runbooks, operational playbooks, escalation procedures, and on-call readiness.
  • Establish platform observability standards covering metrics, logs, traces, business KPIs, and customer experience monitoring.
  • Drive implementation of dashboards that provide executive-level, operational, and application-level visibility.
  • Ensure monitoring solutions provide actionable insights and proactive detection of service degradation.
  • Define alerting strategies, ownership models, escalation paths, and operational response procedures.
  • Partner with engineering teams to continuously improve monitoring coverage and operational intelligence.
  • Lead platform release planning and coordination across multiple delivery teams and Agile Release Trains.
  • Implement standardized release governance, risk management, deployment readiness, and change control processes.
  • Oversee release calendars, dependency management, deployment sequencing, and environment coordination.
  • Ensure compliance with change management requirements and regulatory controls.
  • Monitor release execution metrics and drive improvements in deployment success rates and delivery predictability.
  • Own production deployment governance and operational execution for critical business applications.
  • Establish deployment automation, rollout strategies, rollback procedures, and deployment runbooks.
  • Lead deployment readiness reviews and ensure stakeholder alignment across development, QA, infrastructure, and support teams.
  • Manage production deployment risks and facilitate issue resolution during high-impact releases.
  • Drive continuous improvements that increase deployment frequency while reducing change failure rates.
  • Define operational KPIs including availability, MTTR, deployment success rate, change failure rate, incident volume, and platform health indicators.
  • Lead continuous service improvement initiatives focused on reliability, scalability, resilience, and customer experience.
  • Partner with architecture and engineering teams to influence technology standards and platform modernization efforts.
  • Ensure operational readiness for new applications, services, and infrastructure platforms.
  • Foster a blameless culture focused on learning, accountability, and continuous improvement.

Benefits

  • health and well-being benefits
  • savings and retirement programs
  • paid time off (including Vacation PTO, Flex PTO, and Holiday PTO)
  • banking benefits and discounts
  • career development
  • reward and recognition
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service