Site Reliability Engineer (Hybrid)

Unlimited SystemsCincinnati, OH
Hybrid

About The Position

Unlimited Systems authors the category-leading Unlimited Financials practice management system focused on the unique revenue cycle requirements of specialty healthcare providers. Unlimited Systems customers enjoy streamlined business office workflows, reduced claim denial rates, and accelerated and amplified revenue streams. Unlimited Systems is committed to ensuring that specialty healthcare providers thrive in a dynamic reimbursement environment. Unlimited Systems is a portfolio company of Francisco Partners, a leading technology investment firm with deep sector focus and a track record of delivering outstanding returns. Through private equity and credit funds, they provide flexible capital and partnership to growth-aspiring technology companies. The Challenge Unlimited Financials operates across a sophisticated cloud-native environment — Azure Kubernetes Service, event-sourced microservices, healthcare integrations, and a growing network of downstream data pipelines — all of which requires dedicated technical support to remain stable, observable, and performant for our customers. We are building the reliability practice this platform deserves. As we scale into new specialty healthcare markets and take on greater complexity, we need someone who thinks proactively — not just responding to incidents but engineering the systems and signals that prevent them. This is not a run-and-maintain role. It is a “build-and-shape-the-future" role. Our Site Reliability Engineer will own the observability strategy across our observability applications, define and drive meaningful SLOs and SLIs for our critical services, and work shoulder-to-shoulder with DevOps and Software Engineering to make reliability a first-class concern — not an afterthought. You will bring structure to chaos, clarity to ambiguity, and measurable improvement to customer experience.

Requirements

  • 5+ years in an SRE, DevOps, or Platform Engineering role in a production cloud environment.
  • Hands-on AKS experience: cluster provisioning, upgrades, CNI networking, and workload lifecycle management.
  • Proficiency with Splunk: SPL authoring, dashboard creation, alert configuration, and data onboarding.
  • Working knowledge of Instana APM: agent deployment, custom tracing, alerting, and performance analysis.
  • Solid command of Azure Monitor, Log Analytics (KQL), and Application Insights.
  • Infrastructure-as-Code experience with Terraform; familiarity with Helm and Kustomize.
  • Scripting or development proficiency in at least one of: Python, Go, or Bash.
  • Demonstrated ability to drive incident resolution and lead blameless post-mortems.
  • Clear communicator — able to translate complex technical topics for non-technical stakeholders.

Nice To Haves

  • Microsoft Certified: Azure Administrator Associate (AZ-104) or Azure DevOps Engineer Expert (AZ-400).
  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
  • Splunk Certified Power User or Splunk Enterprise Certified Architect.
  • Experience with service mesh technologies (Istio, Linkerd) on AKS.
  • Familiarity with FinOps practices and Azure cost-management tooling.
  • Experience in a regulated environment (SOC 2, ISO 27001, PCI-DSS, or HIPAA).

Responsibilities

  • Defining, tracking, and reporting on SLIs, SLOs, and error budgets for all critical services.
  • Designing and maintaining runbooks, escalation paths, and on-call rotation schedules.
  • Designing chaos engineering practices to proactively surface reliability weaknesses before they impact customers.
  • Building and maintaining Splunk searches, dashboards, and alert policies covering application and infrastructure logs.
  • Developing KPIs and unified service-health views for engineering and leadership.
  • Configuring and extending instrumentation across microservices for distributed tracing and real-time baselining.
  • Creating smart alerts integrated with on call and ticketing systems for automated incident routing.
  • Maintaining Azure Monitor alert rules, action groups, and workbooks across Azure subscriptions.
  • Utilizing Log Analytics workspaces including data-retention policies and ingestion cost governance using KQL.
  • Leveraging Application Insights for APM, availability testing.
  • Driving convergence of Splunk, Instana, and Azure signals into a unified observability strategy.
  • Building internal tooling in Python, Go, or Bash to eliminate toil and accelerate incident response.
  • Participating in security reviews and threat-modeling sessions for new platform capabilities.

Benefits

  • Unlimited Systems Core Values
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service