Principal Site Reliability Engineer - Remote

UnitedHealth GroupEden Prairie, MN
$134,600 - $230,800Remote

About The Position

Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You'll design systems that help engineers detect issues faster, automate response workflows, improve resiliency, and transform operational data into actionable insights. As a technical leader, you'll influence reliability standards, mentor engineers, and drive next-generation observability and automation strategies across Azure and AWS environments. You’ll enjoy the flexibility to work remotely from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Requirements

  • 10+ years of experience in software engineering, platform engineering, DevOps, or SRE roles
  • 3+ years of experience in a principal, staff, lead, or senior technical leadership role
  • 5+ years of experience with cloud platforms and container orchestration, preferably Azure or AWS
  • 3+ years of experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, or similar platforms
  • 1+ years of experience designing production automation, tooling, or AI-assisted workflows for incident response or operational decision-making

Nice To Haves

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
  • Experience with LLM-based systems, AI agents, RAG, tool orchestration, evaluations, or guardrails
  • Experience with resiliency engineering, disaster recovery, chaos engineering, or recovery validation
  • Experience with infrastructure as code and automation tools such as Terraform, Pulumi, Ansible, Helm, or Kubernetes operators
  • Solid background in incident command, runbooks, postmortems, production readiness, and reliability governance

Responsibilities

  • Build AI-assisted SRE capabilities that accelerate incident detection, triage, mitigation, and recovery
  • Connect observability, deployment, runbook, ownership, and incident data into actionable operational context
  • Design human-in-the-loop workflows for safe mitigation, approvals, recovery verification, and auditability
  • Standardize OpenTelemetry, SLIs, SLOs, error budgets, and reliability scorecards across critical services
  • Improve alert quality by reducing noise, clarifying customer impact, and identifying likely causes faster
  • Lead resiliency testing, DR exercises, chaos engineering, and automated recovery validation
  • Mentor engineers and lead cross-functional reliability improvements across Optum Financial

Benefits

  • comprehensive benefits package
  • incentive and recognition programs
  • equity stock purchase
  • 401k contribution
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service