Senior Site Reliability Engineer

The Semios Group
CA$140,000 - CA$160,000Hybrid

About The Position

As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the scalability, reliability, and performance of our infrastructure and services. Operating within a high-performing engineering team, this role bridges the gap between development and operations, with a strong focus on automation, observability, and resilience. You will use your technical expertise and leadership skills to drive improvements across systems and processes, ensuring we deliver high-quality, reliable products to our customers. This is a hands-on role where your impact will be felt across both technical execution and team development. You will be part of a team operating across time-zones and global regions.

Requirements

  • Have good knowledge of Linux and bash or similar.
  • Be versed in the delivery of a SaaS product on AWS, GCP, or Azure.
  • Have strong programming skills (Ruby, Python, Go, etc.).
  • Be competent with Terraform or similar Infrastructure as Code (IaC) tools.
  • Have experience with Docker, Kubernetes, EKS, or similar technologies.
  • Have experience with CI/CD pipelines on Buildkite or similar platforms.
  • Be familiar with building delivery pipelines with Buildkite or similar.
  • Have strong version control skills with Git.
  • Be experienced with increasing monitoring and observability using Datadog or similar tools (New Relic, Splunk, etc.).
  • Demonstrate a strong automation mindset, with a focus on eliminating repetitive tasks through scripting, tooling, and documentation.
  • Enjoy delivering quickly and iterating fast.
  • 8+ years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering roles supporting production cloud environments.
  • 3+ years of experience in a senior or technical leadership capacity, with demonstrated ownership of critical production systems and mentoring of engineers.
  • Hands-on experience with modern cloud environments (AWS, GCP, or Azure), including deployment, scaling, monitoring, and cost optimization of SaaS applications.
  • 5+ years of relevant experience in DevOps, SRE, or infrastructure engineering roles.
  • Proven experience implementing and managing observability stacks (e.g., Datadog, Prometheus, New Relic, Splunk) and driving improvements to SLIs/SLOs.
  • Experience in incident management, including participation in on-call rotations and leading post-incident reviews with a focus on continuous improvement.

Nice To Haves

  • AI/LLM upskilling — comfort using AI and agentic tooling (e.g., Claude Code) to accelerate investigation, automation, and delivery.
  • Service mesh (Envoy / Istio) — hands-on experience deploying and operating a service mesh for traffic management, observability, and secure service-to-service communication.
  • NATS — experience running or building on NATS (or comparable messaging/streaming systems) for event-driven and distributed architectures.
  • AWS (preferred), GCP/Azure, Terraform, Docker, Kubernetes (EKS), Buildkite/GitHub Actions/Jenkins, Python/Ruby/Go, Datadog/New Relic/Prometheus, Git (GitHub/GitLab), strong Linux.

Responsibilities

  • Lead the delivery of infrastructure projects.
  • Plan and perform higher-risk maintenance.
  • Contribute to resolving incidents and participate in an on-call roster.
  • Work with product and software development colleagues to improve the resiliency and reliability of our products.
  • Mentor team members in all aspects of SRE work.
  • Manage your productivity and workload in a work-from-home environment.
  • Use a data-driven approach to identify changes to the product architecture to improve reliability, performance, and availability.
  • Fully understand production environments and the end-to-end delivery process.
  • Identify parts of the system that do not scale and drive solutions for these problem areas.
  • Maintain and improve Service Level Indicators (SLI) that align with availability and performance targets.
  • Build quality into the team's work by encouraging refactoring, testing, and breaking up the team’s work into small, releasable pieces.
  • Promote automation and continuous improvement to reduce operational overhead and improve platform reliability.

Benefits

  • Generous vacation policy
  • company-paid holidays
  • year-end winter break
  • Hybrid working arrangements
  • comprehensive health plans designed to support your physical and mental health
  • Group RRSP, which includes a 3% company paid match after three months of employment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service