Sr. Manager, Site Reliability

OmnicellAustin, TX
Hybrid

About The Position

Omnicell is establishing a new Global Cloud Operations organization focused on a cloud-native, SaaS-delivered platform. This Site Reliability Engineering role is crucial for building this function from the ground up. The individual will be responsible for designing the SRE practice, setting standards, and initially performing hands-on engineering and incident response. The role involves defining Service Level Objectives (SLOs), incident management processes, observability strategies, and prioritizing reliability investments. The environment is hybrid, operating in a regulated space (HIPAA, SOC 2, FedRAMP) where reliability, security, and auditability are paramount. This position also anchors the company's investment in AI-driven operations, including AIOps and ML-assisted observability, requiring technical ownership of their introduction and validation within the regulated environment.

Requirements

  • Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment, with demonstrated experience building a practice from zero or near-zero.
  • Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.
  • Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, GitHub Actions, Jenkins, TeamCity, or equivalent).
  • Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.
  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations. Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd).
  • Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent.
  • Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be applied in a regulated environment.
  • Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice and documented follow-up discipline.
  • Ability to coach and mentor, with direct evidence of growing junior and mid-level engineers.
  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.
  • Excellent communication and stakeholder management skills; ability to translate complex technical concepts for non-technical audiences.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience.
  • 8+ years of experience in software or platform engineering, with at least 4 of those in an SRE, DevOps, or platform reliability role.
  • Proven Experience advising and influencing senior technical or operations leaders using data driven recommendations.
  • At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.

Nice To Haves

  • Masters Degree
  • Prior experience in healthcare, clinical workflows, or another regulated vertical.
  • Experience transitioning from MSP-heavy operations to internal-first, or integrating managed service providers (IBM, HCL, or similar) into an SRE operating model.
  • Exposure to hybrid hardware-plus-cloud products, where device reliability and cloud reliability are jointly owned.
  • Experience building or integrating AIOps platforms for automated incident triage and remediation.
  • Familiarity with large language model APIs or agentic AI frameworks applied to on-call automation or runbook generation.
  • Experience deploying and managing stateful distributed services in Kubernetes.
  • Hands-on experience with security scanning and intrusion detection systems in regulated environments (HIPAA, SOC 2, or equivalent).
  • Experience with messaging systems such as Kafka or RabbitMQ.
  • Familiarity with chaos engineering principles and tooling (Chaos Monkey, LitmusChaos, or similar).
  • Working knowledge of Databricks, Team Foundation Server, Octopus Deploy, or similar tools in Omnicell's current stack.
  • Experience with FinOps practices and cloud cost optimization strategies.

Responsibilities

  • Define and publish SLOs and SLIs for top Tier-1 customer-facing services, establish error budget policy and enforcement.
  • Design the incident command structure, including severity rubric, declaration criteria, war-room protocol, communication cadence, and postmortem template. Train incident commanders.
  • Select and implement the primary observability platform, defining instrumentation standards.
  • Partner to transition incident response RACI into a durable SRE-owned model.
  • Establish the on-call rotation model, including distribution, compensation, paging discipline, and handoff protocols with managed services partners.
  • Develop and track operational KPIs (MTTR, SLO attainment, change-failure rate, etc.) and present reliability metrics and roadmaps to senior leadership.
  • Instrument Tier-1 services, write dashboards, alerts, and runbooks.
  • Participate in on-call rotation, commanding Sev-1 and Sev-2 incidents, leading blameless postmortems, and driving follow-up work.
  • Contribute code and infrastructure-as-code (Terraform preferred) to the platform.
  • Oversee the design and evolution of CI/CD pipelines.
  • Administer and scale Kubernetes platform, ensuring secure and compliant configurations.
  • Run chaos and failover exercises to validate resilience.
  • Architect Omnicell's AIOps direction, evaluating and introducing ML-based anomaly detection, predictive alerting, automated root cause analysis, and LLM-assisted pipelines.
  • Make build-versus-buy decisions for AIOps tooling and integrate AI-assisted tools into the observability and incident response stack.
  • Ensure AI-assisted operations meet auditability and explainability requirements in a regulated environment.
  • Coach one Engineer III SRE, providing guidance on incidents, design proposals, and professional growth.
  • Design future SRE hires, including writing requisitions and running interview loops.
  • Represent SRE in architecture reviews, product launch readiness reviews, and executive Cloud Ops metric reviews.
  • Partner with Enterprise Security, Compliance, and Architecture to ensure platform services meet regulatory and security requirements.

Benefits

  • On-call participation expected as part of the SRE rotation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service