Team Lead, Site Reliability Engineer (SRE)

Loblaw Companies LimitedBrampton, ON
CA$128,000 - CA$176,000Onsite

About The Position

We're looking for an experienced engineering leader to shape and advance our Site Reliability Engineering practice. In this role, you'll lead a team of talented SREs while driving the strategy, standards, and operational excellence that enable reliable, scalable, and secure platforms across the organization. Balancing technical leadership with people leadership, you will influence engineering direction, guide critical operational decisions, and partner across engineering, infrastructure, security, and product teams to continuously improve reliability, automation, and platform resilience. You will play a key role in building a high-performing team and fostering a culture of ownership, accountability, and continuous improvement.

Requirements

  • Proven experience leading Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering teams in large-scale, production environments.
  • Demonstrated success building, mentoring, and leading high-performing engineering teams while influencing technical strategy and organizational outcomes.
  • Proven ability to influence technical direction, drive strategic initiatives, manage competing priorities, and build strong partnerships across engineering, infrastructure, security, and product teams.
  • Experience defining engineering standards, operational processes, and best practices that improve reliability and scalability across multiple teams.
  • Experience leading major incident response, driving blameless postmortems, and implementing continuous improvements to strengthen operational resilience.
  • Deep expertise in Site Reliability Engineering principles, including service level objectives (SLOs), service level agreements (SLAs), error budgets, capacity planning, and resilient system design.
  • Extensive experience designing, deploying, and operating Kubernetes platforms at scale, including distributed and edge computing environments.
  • Hands-on experience designing and operating cloud infrastructure on AWS, Azure, or GCP, leveraging Infrastructure as Code technologies such as Terraform and configuration management tools like Ansible.
  • Strong knowledge of observability and operational excellence practices, including monitoring, logging, alerting, and distributed tracing using tools such as Prometheus and Grafana.
  • Strong software engineering and automation skills, with proficiency in Python, Bash, Go, or similar scripting languages.
  • Strong understanding of enterprise application ecosystems, including Java-based applications and their operational characteristics.
  • Experience managing and optimizing enterprise database platforms, including DB2, SQL Server, and PostgreSQL.

Responsibilities

  • Drive Reliability Strategy – Define and advance reliability engineering practices that improve availability, scalability, resilience, and operational excellence across our platforms. Establish engineering standards and champion a culture of reliability by design.
  • Influence Engineering Direction – Partner with engineering, infrastructure, security, and product leaders to establish reliability standards, align priorities, and drive strategic initiatives that improve engineering effectiveness across the organization.
  • Build and Lead a High-Performing SRE Team – Mentor, coach, and develop engineers while driving hiring, workforce planning, performance management, succession planning, and career development. Foster a culture of ownership, accountability, collaboration, and continuous learning.
  • Lead Operational Excellence – Own incident management practices, guide major incident response, drive blameless postmortems, and ensure long-term improvements that strengthen platform resilience and operational maturity.
  • Drive Automation Strategy – Champion automation-first principles and invest in tooling that improves developer productivity, reduces operational overhead, and increases platform reliability.
  • Lead Infrastructure Engineering Practices – Establish best practices for Infrastructure as Code, platform provisioning, configuration management, and Kubernetes operations while promoting consistency, security, and scalability.
  • Advance Platform Delivery – Guide the evolution of CI/CD capabilities and deployment practices to enable secure, scalable, and reliable software delivery.

Benefits

  • Work Perks Program
  • On-site Gym, Basketball & Volleyball courts, Ice Rink, Groceries delivered to work via PC Express, Dry Cleaning services (1PCC Office)
  • Tuition Reimbursement & Online Learning
  • Pension & Benefits
  • Paid Vacation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service