Site Reliability Engineer II

NationsBenefits, LLCPlantation, FL
Remote

About The Position

NationsBenefits is a leading Healthcare Fintech provider specializing in supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to deliver innovative healthcare solutions that enhance growth, improve outcomes, reduce costs, and provide value to members. Our comprehensive offerings, including supplemental benefits, fintech payment platforms, and member engagement tools, empower health plans to offer high-quality benefits that address social determinants of health, thereby improving member health outcomes and satisfaction. Our robust, compliance-focused infrastructure, proprietary technology, and premier service model enable us to support health plans in delivering value-based care to millions of members. We foster a rewarding work environment that attracts top talent and encourages contributions to premier service for both internal and external customers, aiming to positively transform the healthcare industry. We provide internal career advancement opportunities across our US, South America, and India locations.

Requirements

  • 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support.
  • Hands-on experience with production incident response, troubleshooting, and escalation.
  • Experience with Datadog or similar monitoring and observability platforms.
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management.
  • Experience with Docker or other container technologies.
  • Working knowledge of SQL, MySQL, or NoSQL databases.
  • Ability to work effectively in high-volume, mission-critical production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent written and verbal communication skills.
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model.

Nice To Haves

  • Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP).
  • Familiarity with CI/CD pipelines and deployment automation.
  • Experience with Helm Charts and Kubernetes deployments.
  • Knowledge of ITIL principles and Agile methodologies.
  • Experience supporting regulated environments such as healthcare or fintech.
  • Scripting or programming experience in Python, PowerShell, Bash, Java, or C#.

Responsibilities

  • Serve as the first responder for production incidents by identifying, triaging, and resolving issues.
  • Monitor and respond to alerts generated by Datadog and other monitoring platforms.
  • Perform initial root cause analysis and escalate incidents according to defined SLAs.
  • Communicate incident status and resolution updates to internal stakeholders.
  • Partner with senior engineers to resolve complex production issues.
  • Continuously monitor application health, infrastructure performance, and system availability.
  • Configure and optimize monitoring dashboards and alert thresholds.
  • Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis.
  • Support containerized applications running in Kubernetes and Docker environments.
  • Participate in a weekday "Follow-the-Sun" production support model with global engineering teams.
  • Participate in an on-call rotation for critical production systems as needed.
  • Help maintain high availability and system uptime.
  • Develop automation scripts and operational tools using one or more of the following: Python, PowerShell, Bash, C#, Java.
  • Support CI/CD pipeline monitoring and deployment reliability.
  • Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks.
  • Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams.
  • Recommend improvements to monitoring, tooling, and operational processes.
  • Collaborate effectively with globally distributed engineering teams.
  • Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews.
  • Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.

Benefits

  • Competitive compensation
  • Comprehensive benefits
  • Unlimited Paid Time Off (PTO)
  • Opportunities for career growth and professional development
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service