Lead Site Reliability Engineer

RBCToronto, ON
Onsite

About The Position

Join our Credit Technology team as a Lead Site Reliability Engineer, where you'll play a key role in driving operational excellence through technology, process optimization, and cross-functional collaboration for the Personal Credit SRE & Ops team. This exciting opportunity will challenge you to work with cutting-edge technologies, including AI and emerging innovations, and collaborate closely with development teams to deliver embedded SRE solutions. As a vital link between QE, DevOps, Development, Infrastructure, and Support teams, you'll leverage your strong technical skills to solve complex problems and drive success across multiple components and technologies. If you're passionate about tackling new challenges and developing innovative solutions, we invite you to join our team and take your career to the next level.

Requirements

  • 5-7 years of experience as a Site Reliability Engineer or a Cloud Developer
  • Proven leadership experience managing cross-functional teams and stakeholders
  • Good understanding of Kubernetes and Cloud working knowledge with experience and understanding of CICD pipeline and DevOps / Agile Methodology
  • Decent knowledge of the following SRE practices and technologies: Python, YAML, Shell scripting, OpenShift, Linux, MongoDB, Dynatrace, Prometheus, PagerDuty, Moog, Splunk, Elastic, Ansible, Grafana, Chaos Engineering, MQ, Kafka
  • Perform production support role, including off-hours support
  • Excellent communication skills

Nice To Haves

  • Experience in the context of SRE and/or Application Development, Test Automation teams

Responsibilities

  • Act as one of the final escalation points for critical outages and lead 24/7 incident response with rapid resolution of customer-impacting issues
  • Lead strategic direction and continuous improvement initiatives across Credit Technology
  • Manage cross-functional teams and stakeholders to execute upgrade and operational change management
  • Oversee end-to-end reliability of the ecosystem (hardware, software, network) ensuring 99.9% availability
  • Generate performance metrics (SLOs/SLAs/SLIs) and maintain regulatory compliance and security standards
  • Serve as primary relationship owner for vendor services, maintenance, and internal technology teams
  • Identify, design, write and test automation procedures using AI, Ansible and other relevant technologies
  • Support applications running on many platforms including OpenShift and distributed systems
  • Implement Chaos Engineering experiments and Disaster Recovery procedures to test and validate system resilience and reliability
  • Implement monitoring and alerting, anomaly detection and reliability testing for applications in scope

Benefits

  • bonuses
  • flexible benefits
  • competitive compensation
  • commissions
  • stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive and high-performing team
  • A world-class training program in financial services
  • Flexible work/life balance options
  • Opportunities to do challenging work
  • Opportunities to take on progressively greater accountabilities
  • Opportunities to building close relationships with clients
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service