Site Reliability Engineering Manager

NationsBenefitsPlantation, FL
Remote

About The Position

NationsBenefits is a rapidly growing company in the Healthcare Fintech sector, providing supplemental benefits, flex cards, and member engagement solutions. They partner with managed care organizations to offer innovative healthcare solutions that enhance growth, improve outcomes, reduce costs, and deliver value to members. Their offerings include supplemental benefits, fintech payment platforms, and member engagement solutions designed to help health plans provide high-quality benefits that address social determinants of health, thereby improving member health outcomes and satisfaction. The company's compliance-focused infrastructure, proprietary technology, and premier service model enable health plan partners to deliver value-based care to millions of members. NationsBenefits fosters a fulfilling work environment that attracts top talent and encourages contributions to premier service for both internal and external customers, aiming to positively transform the healthcare industry. They offer career advancement opportunities across multiple US, South America, and India locations.

Requirements

  • 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • 1–2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player-coach leadership model.
  • Strong hands-on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands-on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
  • Experience with Helm, CI/CD pipelines, and deployment automation.
  • Working knowledge of ITIL processes and Agile methodologies.
  • Experience working with SQL, MySQL, or NoSQL databases.
  • Excellent communication and stakeholder management skills.
  • Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.

Nice To Haves

  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience building or scaling SRE teams and on-call programs.
  • Experience defining and managing SLOs, SLIs, and error budgets.
  • Prior experience in the healthcare or fintech industry.
  • Knowledge of security and compliance frameworks relevant to regulated environments.

Responsibilities

  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.
  • Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry PagerDuty on-call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self-healing capabilities, and runbook maturity.
  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
  • Contribute hands-on to tooling, automation, and technical reviews as needed.
  • Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
  • Represent the US SRE organization in cross-functional planning and operational reviews.
  • Communicate effectively with both technical and non-technical stakeholders.
  • Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.

Benefits

  • Competitive compensation
  • Comprehensive benefits
  • Unlimited PTO
  • Fully remote work environment (US-based)
  • Opportunity to lead and grow a high-impact SRE organization
  • Exposure to modern cloud-native technologies and large-scale reliability challenges
  • Collaborative culture focused on innovation, learning, and continuous improvement
  • Meaningful work that directly impacts healthcare technology and millions of members
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service