Senior Principal Site Reliability Engineer

Questrade Financial GroupToronto, ON
CA$150,000 - CA$190,000Hybrid

About The Position

Questrade Financial Group (QFG) is seeking a Senior Principal Site Reliability Engineer to ensure the stability, resiliency, and scalability of critical brokerage back-end applications across a hybrid on-premises and cloud architecture. This role involves driving reliability engineering practices such as SLOs/SLIs, observability, incident response, and capacity planning. The engineer will also make hands-on code contributions using an inner-source model, identify systemic reliability risks, and improve operational standards for various teams. This position is ideal for a senior principal engineer who enjoys fixing production reliability at its source, is comfortable with multiple codebases and environments, and aims to significantly impact the resiliency of a regulated, high-availability brokerage platform.

Requirements

  • Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent combination of education and experience.
  • 8+ years of software engineering and/or site reliability engineering experience, including production ownership of business-critical applications; financial services or brokerage experience strongly preferred.
  • Demonstrated ability to read, debug, and make minor-to-moderate code changes across multiple languages/stacks (e.g., Java, .NET, Node.js/TypeScript, Python) in an inner-source or cross-team contribution model.
  • Deep experience with cloud scalability strategies on one or more major providers (AWS, Azure, GCP), including auto-scaling, load balancing, multi-region resiliency, and cost-aware capacity planning.
  • Experience operating and supporting hybrid architectures spanning on-premises data centers and cloud environments.
  • Strong background in observability tooling (e.g., Prometheus/Grafana, Datadog, Splunk, ELK, AppDynamics, Dynatrace) and building actionable alerting and dashboards.
  • Practical experience defining and operating against SLOs/SLIs/error budgets and running blameless post-incident reviews.
  • Experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, Ansible, CloudFormation) in support of reliable, repeatable deployments.
  • Solid understanding of microservices architecture, distributed systems failure modes, and resiliency patterns (circuit breakers, retries/backoff, bulkheads, timeouts).
  • Familiarity with relational and NoSQL data stores and their operational/scaling characteristics.
  • Experience with incident management and on-call practices (e.g., PagerDuty, Opsgenie) including leading major incident response
  • Knowledge of security, audit, and regulatory considerations relevant to brokerage / financial services production systems.
  • Excellent communication skills, with the ability to influence engineers and stakeholders across many teams without direct authority.
  • Strong documentation, analytical, and problem-solving skills

Responsibilities

  • Own the end-to-end reliability posture of critical brokerage back-end applications, driving measurable improvements in availability, latency, and error budgets across on-premises and cloud environments.
  • Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets in partnership with application teams; use them to prioritize reliability work over feature work when warranted.
  • Lead root cause analysis and blameless post-incident reviews for high-severity production incidents; drive remediation items to closure and identify systemic patterns across applications.
  • Establish and mature observability practices (metrics, logging, tracing, alerting) so that failures are detected proactively and diagnosed quickly across a heterogeneous, multi-stack estate.
  • Build capacity planning, load testing, and chaos/failure-injection practices to validate resilience before incidents occur.
  • Champion a culture of operational excellence, toil reduction, and AI and automation-first thinking across engineering teams.
  • Make hands-on, code contributions directly into multiple applications spanning different languages, frameworks, and stacks, using an inner-source model to fix reliability defects, add instrumentation, and improve resiliency patterns.
  • Partner with individual application teams to raise pull requests, follow their contribution standards, and pair with owning engineers so fixes land safely and are properly reviewed and owned long-term.
  • Identify recurring reliability anti-patterns across codebases (e.g., missing timeouts/retries, unbounded queues, improper connection pooling) and drive standardized, reusable fixes or shared libraries.
  • Contribute to and help govern internal reliability tooling, shared SDKs, and common patterns (circuit breakers, backoff/retry, health checks) that can be inner-sourced across teams.
  • Design and advise on cloud scalability strategies (auto-scaling, load balancing, multi-region/multi-AZ failover, caching, queuing) for workloads that span on-premises data centers and public cloud.
  • Guide capacity and cost-aware scaling decisions, balancing performance, resiliency, and cloud spend across hybrid deployments.
  • Evaluate and recommend cloud-native and hybrid resiliency patterns (e.g., disaster recovery, active-active/active-passive architectures, data replication strategies) appropriate for regulated brokerage workloads.
  • Bring strong organizational awareness of the operational, financial, regulatory, and reputational risk that production incidents pose to a brokerage business, and factor that into prioritization.
  • Participate in risk assessments related to system reliability, availability, and disaster recovery, partnering with Risk, Compliance, and Information Security as needed.
  • Contribute to change management and release governance practices that reduce the likelihood and blast radius of production incidents.
  • Promptly identify, escalate, and help remediate reliability or security-related incidents in accordance with company policy.
  • Act as a technical reference and mentor for reliability engineering practices, coaching application teams on operational excellence without formal direct reports.
  • Influence architecture and design decisions across multiple teams by bringing a reliability and scalability lens to reviews and planning.
  • Document and evangelize reliability standards, runbooks, and best practices; lead or contribute to internal tech talks and communities of practice.
  • Partner with engineering leadership to define the reliability roadmap and report on progress against stability goals.

Benefits

  • Health & wellbeing resources and programs
  • Paid vacation, personal, and sick days for work-life balance
  • Competitive compensation and benefits packages
  • Career growth and development opportunities
  • Opportunities to contribute to community causes
  • Comprehensive benefits plan
  • Competitive incentive (bonus) program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service