About The Position

Manager – Production Operations & Site Reliability Engineering At Alcon, we are driven by the meaningful work we do to help people see brilliantly. We innovate boldly, champion progress, and act with speed as the global leader in eye care. Here, you’ll be recognized for your commitment and contributions and see your career like never before. Together, we go above and beyond to make an impact in the lives of our patients and customers. We foster an inclusive culture and are looking for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability, availability, security, and continuous improvement of Alcon’s Digital Health Cloud platform supporting Software as a Medical Device (SaMD) and customer-facing digital health applications. This role serves as the senior technical authority for production operations and Site Reliability Engineering (SRE), driving platform reliability, observability, automation, incident management, release governance, and cloud optimization. Partnering across Product Engineering, Architecture, Security, Infrastructure, and Operations, the Principal Engineer enables resilient, compliant, and scalable healthcare platforms with predictable, high-quality software delivery.

Requirements

  • Bachelor’s Degree or Equivalent years of directly related experience (or high school +13 yrs; Assoc.+9 yrs; M.S.+2 yrs; PhD+0 yrs)
  • The ability to fluently read, write, understand and communicate in English
  • 5 Years of Relevant Experience

Nice To Haves

  • Bachelor’s Degree in degree in Computer Science, Engineering, or related field (Master’s preferred)
  • Experience in Production Operations, Site Reliability Engineering, Platform Engineering, or Cloud Operations
  • Experience supporting regulated healthcare or medical device platforms
  • AWS Solutions Architect or DevOps Professional certification
  • Kubernetes (CKA/CKS) certification
  • Experience with healthcare interoperability standards (HL7, FHIR, DICOM)
  • Experience implementing Site Reliability Engineering practices within large-scale cloud environments
  • AWS (EKS, EC2, RDS, S3, Route53, IAM, Load Balancers, CloudWatch)
  • Kubernetes, Istio, Docker
  • Datadog, APM, logging, distributed tracing, synthetic monitoring
  • CI/CD, Infrastructure as Code, automation
  • HIPAA, GDPR, FDA compliance
  • Strong understanding of cloud networking, security, and high-availability architectures

Responsibilities

  • Lead steady-state operations for cloud-based healthcare platforms, ensuring high availability, reliability, and performance.
  • Establish and continuously improve operational standards, SLIs/SLOs, readiness reviews, and service excellence practices.
  • Drive platform resilience, capacity planning, disaster recovery, and governance through Alcon’s Steady State Operations Framework (SSOF).
  • Lead production readiness reviews, ensuring applications meet operational, security, monitoring, compliance, and supportability requirements.
  • Govern the Road to Production process across Validation, Staging, and Production environments, validating deployment readiness, infrastructure qualification, rollback plans, and acceptance criteria.
  • Partner with Engineering and DevOps teams to improve release quality, deployment reliability, and change success through standardized governance and automation.
  • Champion SRE best practices, leveraging automation and self-healing capabilities to reduce operational toil.
  • Improve service reliability and customer experience by reducing MTTD and MTTR and increasing deployment success rates.
  • Lead incident investigations, root cause analyses, and long-term corrective actions for critical production events.
  • Provide technical leadership across AWS platforms including EKS, EC2, RDS, S3, ElastiCache, AWS MQ, Route53, and Kubernetes/Istio.
  • Optimize cloud infrastructure for scalability, resilience, security, performance, and cost efficiency.
  • Define enterprise observability strategies using Datadog, CloudWatch, distributed tracing, synthetic monitoring, centralized logging, and executive dashboards.
  • Lead automation initiatives across deployments, monitoring, health validation, incident response, and operational workflows.
  • Ensure compliance with HIPAA, GDPR, FDA, and enterprise cybersecurity standards.
  • Partner with Security teams to strengthen cloud security architecture, identity management, network segmentation, and operational controls.
  • Serve as the senior escalation point for major incidents, production events, and Hypercare operations.
  • Mentor engineering teams on operational excellence, production engineering, and SRE best practices.
  • Influence platform and architectural decisions that enhance operability, maintainability, resilience, and long-term service reliability.
  • Collaborate across Engineering, Architecture, Infrastructure, Security, and Global Operations to advance platform stability and operational maturity.

Benefits

  • health
  • life
  • retirement
  • PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service