Sr. Manager, Site Reliability

OmnicellAustin, TX
Hybrid

About The Position

Omnicell is establishing a new Global Cloud Operations organization to support its transition from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform. The Site Reliability Engineering (SRE) function is central to this effort. This role is the inaugural senior SRE position, responsible for designing the SRE practice, setting reliability standards, and initially executing these responsibilities directly until the team is sufficiently staffed. The position requires defining what 'good' looks like for Omnicell in terms of service SLOs, incident management, on-call procedures, observability platforms, and prioritizing reliability investments. The role operates in a hybrid environment with both cloud services and on-premise hardware, serving customers through private circuits and the public internet. Operating within a regulated environment (HIPAA, SOC 2, FedRAMP), reliability, security, and auditability are critical and interconnected concerns. The successful candidate will be comfortable with this complexity and will help the organization design for it. Furthermore, this role anchors Omnicell's investment in AI-driven operations, with plans to integrate AIOps and ML-assisted observability (anomaly detection, alert correlation, LLM-assisted runbooks) within the first year. The SRE will be the technical owner of this integration, balancing it with foundational reliability work and ensuring compliance within the regulated environment.

Requirements

  • Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment — with demonstrated experience building a practice from zero or near-zero: you have set SLOs, defined incident command, and introduced error budget thinking to an organization that did not have it.
  • Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.
  • Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, GitHub Actions, Jenkins, TeamCity, or equivalent).
  • Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.
  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations. Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd).
  • Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like.
  • Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be applied in a regulated environment.
  • Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice and documented follow-up discipline.
  • Ability to coach and mentor, with direct evidence of growing junior and mid-level engineers.
  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.
  • Excellent communication and stakeholder management skills; ability to translate complex technical concepts for non-technical audiences.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience
  • 8+ years of experience in software or platform engineering, with at least 4 of those in an SRE, DevOps, or platform reliability role.
  • Proven Experience advising and influencing senior technical or operations leaders using data driven recommendations.
  • At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.

Nice To Haves

  • Masters Degree
  • Prior experience in healthcare, clinical workflows, or another regulated vertical.
  • Experience transitioning from MSP-heavy operations to internal-first, or integrating managed service providers (IBM, HCL, or similar) into an SRE operating model.
  • Exposure to hybrid hardware-plus-cloud products, where device reliability and cloud reliability are jointly owned.
  • Experience building or integrating AIOps platforms for automated incident triage and remediation.
  • Familiarity with large language model APIs or agentic AI frameworks applied to on-call automation or runbook generation.
  • Experience deploying and managing stateful distributed services in Kubernetes.
  • Hands-on experience with security scanning and intrusion detection systems in regulated environments (HIPAA, SOC 2, or equivalent).
  • Experience with messaging systems such as Kafka or RabbitMQ.
  • Familiarity with chaos engineering principles and tooling (Chaos Monkey, LitmusChaos, or similar).
  • Working knowledge of Databricks, Team Foundation Server, Octopus Deploy, or similar tools in Omnicell's current stack.
  • Experience with FinOps practices and cloud cost optimization strategies.

Responsibilities

  • Define and publish SLOs and SLIs for the top 5–10 Tier-1 customer-facing services, in partnership with Product and Engineering.
  • Establish error budget policy and the enforcement mechanism when budgets burn.
  • Design the incident command structure: severity rubric, declaration criteria, war-room protocol, stakeholder communication cadence, and the postmortem template.
  • Train the first cohort of incident commanders across Engineering and Support.
  • Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement.
  • Define the instrumentation standards all new services must meet.
  • Partner with the VP to migrate the interim incident response RACI into a durable SRE-owned model.
  • Establish the on-call rotation model, including fair distribution, compensation approach, paging discipline, and the handoff protocol with managed services partners.
  • Develop and track operational KPIs — MTTR, SLO attainment, change-failure rate, recurrence, cost per workload — and present reliability metrics and improvement roadmaps to senior leadership.
  • Instrument Tier-1 services yourself, write dashboards, alerts, and runbooks.
  • Take the pager and commander Sev-1 and Sev-2 incidents until a broader on-call rotation is staffed.
  • Lead blameless postmortems and drive follow-up work to resolution.
  • Contribute code and infrastructure-as-code (Terraform preferred) to the platform.
  • Oversee the design and evolution of CI/CD pipelines.
  • Administer and scale Kubernetes platform, including secure and compliant cluster configurations.
  • Run chaos and failover exercises to validate resilience.
  • Architect Omnicell's AIOps direction: evaluate and introduce ML-based anomaly detection, predictive alerting, automated root cause analysis, and LLM-assisted runbook or triage pipelines.
  • Make informed build-versus-buy calls across the AIOps landscape.
  • Integrate AI-assisted tooling into the observability and incident response stack where it adds measurable value.
  • Ensure AI-assisted operations meet the auditability and explainability bar required in a HIPAA and SOC 2 environment.
  • Coach one Engineer III SRE who joins shortly after you do, including pairing on incidents and reviewing design proposals.
  • Design the next 2–4 SRE hires, write requisitions, run interview loops, and make hiring decisions.
  • Represent SRE in architecture reviews, product launch readiness reviews, and executive metric reviews.
  • Partner with Enterprise Security, Compliance, and Architecture to ensure platform services meet regulatory and security requirements.

Benefits

  • On-call participation expected as part of the SRE rotation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service