Senior Manager, Site Reliability Engineer - Remote

UnitedHealth GroupBasking Ridge, NJ
$112,700 - $193,200Remote

About The Position

Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care’s most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and Operational Excellence initiatives across Optum bank. This leader will be responsible for improving service reliability, operational resiliency, deployment automation, observability, incident management, and production readiness for critical banking platforms. The ideal candidate combines strong technical expertise with operational leadership experience, driving engineering excellence, automation, reliability, and continuous improvement while ensuring technology services meet business, customer, regulatory, and operational expectations. This role will also help identify and implement emerging automation and AI-enabled operational capabilities that improve service health, reduce operational toil, and accelerate engineering productivity.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field
  • 10+ years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or Technology Operations
  • 5+ years of experience leading engineering or operational teams
  • Proven experience supporting large-scale, business-critical production environments
  • Experience with SRE principles and practices
  • Experience with DevOps and CI/CD
  • Experience with ITSM processes
  • Experience with Cloud platforms (Azure, AWS)
  • Experience with On-prem environments
  • Experience with Infrastructure-as-Code
  • Experience with Container platforms (Kubernetes, OpenShift)
  • Experience with Observability and monitoring platforms
  • Experience with Incident Management
  • Experience with Problem Management
  • Experience with Change Management
  • Experience with Disaster Recovery
  • Experience with Business Continuity
  • Experience with Service Reliability Programs
  • Experience leading operational transformations and continuous-improvement initiatives

Nice To Haves

  • Experience in banking, financial services, healthcare, or other highly regulated industries
  • Experience implementing enterprise observability solutions such as Datadog, Splunk, Grafana, Prometheus, or OpenTelemetry
  • Experience with cloud-native architectures and platform engineering practices
  • Experience deploying AIOps, ChatOps, or intelligent automation solutions
  • Familiarity with Agentic AI workflows
  • Familiarity with LLM-powered operational tooling
  • Familiarity with Knowledge management platforms
  • Familiarity with AI-enabled incident management solutions
  • Experience establishing SLO frameworks and reliability governance programs

Responsibilities

  • Lead and develop multidisciplinary teams responsible for Site Reliability Engineering, DevOps, Platform Engineering, ITSM, and Operational Excellence
  • Establish and execute enterprise reliability, availability, resiliency, and operational maturity strategies
  • Drive engineering excellence through automation, observability, operational readiness, and continuous improvement practices
  • Partner with Technology, Operations, Security, Infrastructure, Risk, and Business leaders to improve service reliability and customer experience
  • Build and mentor high-performing teams while fostering accountability, innovation, operational ownership, and learning
  • Manage staffing, capacity planning, talent development, succession planning, and organizational growth
  • Establish operational metrics, governance standards, and service review processes to improve service performance and risk management
  • Lead enterprise SRE practices including SLI/SLO adoption, error-budget management, reliability engineering, and operational maturity assessments
  • Drive DevOps transformation initiatives, emphasizing automation, deployment standardization, CI/CD pipelines, Infrastructure-as-Code, and GitOps practices
  • Establish production readiness standards and operational acceptance criteria for new technology deployments
  • Improve platform resiliency through capacity planning, disaster recovery, fault tolerance, and resilience testing
  • Drive reduction of operational toil through automation and self-healing capabilities
  • Partner with application and infrastructure teams to improve system scalability, availability, and performance
  • Lead initiatives to improve deployment frequency, reduce change failure rates, and accelerate service recovery times
  • Establish and mature Incident, Problem, Change, Release, and Service Request Management processes
  • Lead major incident management programs and executive communications during critical service disruptions
  • Drive root-cause analysis and problem-management practices to eliminate recurring incidents
  • Improve operational scorecards, service health reviews, and reliability reporting for executive stakeholders
  • Ensure compliance with regulatory, audit, risk, and operational governance requirements
  • Partner with Technology and Business leaders to improve service quality and customer outcomes through data-driven operational improvements
  • Champion a culture of operational excellence and continuous service improvement
  • Identify opportunities to leverage AI and automation to improve operational effectiveness and engineering productivity
  • Lead implementation and evaluation of solutions involving: AIOps, Intelligent alert correlation, Automated incident triage , Root cause analysis assistance, Knowledge management copilots, Agentic operational workflows etc.
  • Partner with enterprise AI teams to evaluate emerging technologies that improve reliability and operational efficiency
  • Drive responsible adoption of AI-enabled engineering and operational practices
  • Support proof-of-concept initiatives that demonstrate measurable reductions in operational effort and incident resolution times
  • Collaborate with Engineering, Infrastructure, Security, Architecture, Risk, Compliance, and Operations teams to prioritize reliability and operational improvements
  • Serve as a trusted advisor on reliability engineering, operational excellence, and automation strategies
  • Drive alignment between technology and business stakeholders to improve service quality and operational outcomes
  • Influence technology investment decisions that improve platform stability, resiliency, and operational efficiency

Benefits

  • comprehensive benefits package
  • incentive and recognition programs
  • equity stock purchase
  • 401k contribution
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service