Staff Site Reliability Engineer

Lightspeed Commerce, Inc.Montreal, QC
Remote

About The Position

We’re looking for a Staff Site Reliability Engineer to join our Data team in Canada. As a Staff Data SRE, you are the technical backbone of the Data Office's infrastructure platform. Your scope spans the entire Data Business Unit — you solve systemic, platform-level problems, not individual tickets. You own the reliability, scalability, and developer experience of the data platform, and you act as a force multiplier: making the engineers around you faster, the systems more resilient, and the platform easier to consume. You bring deep GCP expertise and a product mindset to data infrastructure. You are comfortable navigating ML/AI workload requirements (Vertex AI, feature stores, training pipelines) and can reason confidently about the BI tooling and data layers that feed it (Looker, BigQuery). Your decisions carry weight across teams, and you communicate them clearly to both engineers and non-technical stakeholders.

Requirements

  • Deep expertise across GCP compute, networking, IAM, GKE, data services, and FinOps.
  • Terraform as primary tool for IaC Proficiency.
  • Hands-on experience with Looker infrastructure and ML/AI platform tooling (Vertex AI, model serving, training pipelines).
  • Proficient in Bash and Golang; Python a plus for data tooling.
  • Strong experience with metrics, logs, traces, alerting, and SLO/SLI design, DataDog,...
  • Product Mindset: Focus on making the platform easy to use, not just 'available.'
  • Experience with Github actions, Circle CI, GCP Cloud Build, …
  • Ability to communicate infrastructure trade-offs to both engineers and non-technical stakeholders.
  • AI proficiency, Go-to AI expertise within the Data BU — evaluates tooling, drives cross-team AI-first development practices.
  • Security-First Mindset, every design decision evaluated through a security and compliance lens.
  • Self-awareness with a willingness to learn and improve.
  • Ability to mentor and train other Engineers in the team.

Responsibilities

  • Own and be accountable for advancing the Data Office infrastructure engineering practices and delivering high-leverage platform projects.
  • Design and deliver significant infrastructure improvements.
  • Write production-quality IaC (Terraform).
  • Lead solutions from design through delivery.
  • Address the hardest platform problems.
  • Manage data infrastructure for batch/streaming workloads, ML/AI environments (Vertex AI, model serving, GPU-backed compute), and the BI serving layer (Looker infrastructure and GCP integration).
  • Lead architecture and design discussions for larger cross-team projects.
  • Review critical PRs.
  • Drive solution design sessions.
  • Provide technical direction that keeps the team coherent and moving forward.
  • Lead conversations in agile ceremonies, incident reviews, and cross-functional syncs.
  • Actively mentor Senior and Intermediate SREs.
  • Set the bar for IaC quality, observability practices, and operational discipline through reviews, pairing, and documentation.
  • Proactively identify risks in complex data migrations or infrastructure changes and propose mitigation strategies to ensure zero data loss and minimal downtime.
  • Participate in on-call rotation and incident response.
  • Contribute as part of the wider team to achieve organization-wide objectives.

Benefits

  • Flexible work environment
  • Culture that celebrates performance
  • Career-defining opportunities
  • Flexible paid time off and remote work policies
  • Equity options
  • Contributions to your pension plan
  • Training opportunities to grow your skills and career
  • Health and wellness credit
  • Time off to volunteer and give back to your community
  • Interest groups, employee led networks, social committees to sponsored sports teams
  • Computer purchase program to get your personal Macbook
  • Enhanced parental leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service