Sr Site Reliability Engineer

HealthEquity•Remote,
•$120,500 - $157,000•Remote

About The Position

As a Senior Site Reliability Engineer at HealthEquity, you'll play a critical role in ensuring the platforms our members depend on remain reliable, resilient, and secure. Millions of people trust HealthEquity to manage their HSAs, FSAs, HRAs, and other benefits — and the systems behind that experience need to be available when members need them most, whether that's during open enrollment, tax season, or a routine claim submission. You'll own operational readiness for key services, drive reliability initiatives that reduce risk and toil across teams, and serve as a technical leader whose decisions directly shape the stability and scalability of our platforms.

Requirements

  • 8+ years of experience with complex production ownership in SRE, infrastructure engineering, or related disciplines
  • Advanced proficiency in distributed systems troubleshooting and performance tuning
  • Experience with cloud infrastructure platforms and tooling (Azure, Terraform, Kubernetes/AKS, or equivalent)
  • Experience with CI/CD pipelines (Azure DevOps, GitHub Actions, or equivalent)
  • Experience with observability, monitoring, and incident management tools (Dynatrace, PagerDuty, or equivalent)
  • Experience applying AI/ML tools to improve debugging, automation, or operational insights (e.g., anomaly detection, automated root cause analysis, AI-assisted coding tools)
  • Strong coding/scripting skills in Python, PowerShell, Go, Bash, or similar languages, with the ability to mentor others in software reliability best practices
  • Demonstrated ability to independently own service-level reliability decisions, incident management, and operational excellence
  • Deep expertise in cloud-native architectures, high-availability systems, and fault-tolerant design patterns
  • Experience analyzing complex systems holistically and making data-driven decisions to improve uptime and performance

Nice To Haves

  • Degree in Computer Science, Information Technology, Engineering, or related field a plus, but not required

Responsibilities

  • Leading reliability improvements for critical services and influencing adjacent teams on production best practices
  • Independently driving complex reliability initiatives, including root cause resolution and long-term corrective actions
  • Designing resilient, scalable systems with appropriate failure handling, observability, and recovery strategies
  • Actively participating in architecture reviews to ensure operational scalability and service resilience
  • Owning service-level operational readiness, SLOs, and incident quality for key platforms or services
  • Strengthening operational controls while balancing delivery speed and risk
  • Applying AI to improve debugging, automation, and operational insights — while modeling effective and responsible use for others
  • Mentoring engineers and helping elevate team capability in reliability engineering practices
  • Influencing team practices and coordinating effectively during incidents and cross-functional operational events
  • Participating in on-call as needed, owning production health and SLOs, and improving reliability and MTTR across teams

Benefits

  • Medical, dental, and vision
  • HSA contribution and match
  • Dependent care FSA match
  • Uncapped paid time off
  • Paid parental leave
  • 401(k) match
  • Personal and healthcare financial literacy programs
  • Ongoing education & tuition assistance
  • Gym and fitness reimbursement
  • Wellness program incentives
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service