Lead Site Reliability Administrator

Open Text CorporationMississauga, ON
CA$118,000 - CA$177,000

About The Position

Opentext is a leading innovator in cloud solutions, dedicated to providing robust and scalable infrastructure to support our clients' needs. We are seeking a talented Site Reliability Engineer (SRE) to join our dynamic team and help us maintain and improve our cloud operations.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • Azure DevOps certification or equivalent certifications.
  • Familiarity with observability tools such as Data Dog, Zabbix, Grafana, ELK stack, or AI-enhanced monitoring platforms.
  • Proficiency in scripting languages such as PowerShell, Python, Bash, or Go.
  • Experience presenting reliability strategies to leadership and driving cross-functional initiatives.
  • Exposure to AI/ML-based monitoring, anomaly detection, or automated remediation systems.
  • Interest in contributing to internal AI adoption strategies for cloud operations.

Responsibilities

  • Architect and manage Kubernetes clusters supporting pods across multi-region deployments to ensure high availability and scalability.
  • Design and maintain Azure -based infrastructure with a focus on security, performance, and cost optimization.
  • Develop and maintain Helm charts for consistent and automated application deployment.
  • Use Terraform to provision and manage infrastructure as code across multiple environments.
  • Monitor and optimize Windows and Linux-based systems, ensuring performance, reliability, and compliance with operational standards.
  • Collaborate with development teams to ensure seamless CI/CD integration and deployment of microservices.
  • Implement and maintain CI/CD pipelines using Octopus Deploy, Jenkins, GitLab CI, and GitHub Actions.
  • Lead root cause analysis and resolution of infrastructure and application performance issues.
  • Participate in on-call rotations and lead incident response for critical systems, ensuring rapid recovery and postmortem analysis.
  • Leverage generative AI tools to accelerate scripting, documentation, troubleshooting, and automation tasks.
  • Explore and contribute to AI-driven observability, alerting, and self-healing strategies for cloud infrastructure and applications, including anomaly detection and automated remediation.

Benefits

  • vacation entitlement
  • paid time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service