Senior SRE/AIOps Engineer

RBCToronto, ON
Onsite

About The Position

This role is responsible for designing, implementing, and maintaining SRE (Site Reliability Engineering) and AIOps (Artificial Intelligence for IT Operations) capabilities to ensure system reliability, proactive monitoring, and automation of self-healing operations. In addition to day-to-day support, the position provides end-to-end operational ownership across systems managed by multiple enterprise teams, including incident coordination, dependency management, and escalation. The role is also responsible for key security and compliance functions such as service ID and certificate management, SSO updates, vulnerability remediation, and lifecycle management of end-of-life components. Our team supports a portfolio of multi-platform HR data pipelines that move and process data into Snowflake through multiple integrated components, requiring end-to-end monitoring, coordination, and support across systems. In addition, we support SaaS-based applications that are primarily vendor-managed, while we retain responsibility for integration, access management, monitoring, and operational oversight.

Requirements

  • 3+ years of SRE or Systems Engineering experience with strong technical expertise.
  • Experience with ServiceNow, ITSM processes including incident, problem, change, and release management.
  • Demonstrated ability to work independently, take ownership, and drive projects to completion.
  • Knowledge of containerized platforms such as Kubernetes or OpenShift.
  • Experience working with Ansible Automation Platform or strong willingness to learn.
  • In-depth knowledge of monitoring and observability tools such as Dynatrace, Elasticsearch, Moogsoft, Catchpoint, and PagerDuty.
  • Knowledge and experience with scripting languages such as BASH, Python and PowerShell.
  • Experience with Linux and Windows Server administration.
  • Experience with or strong interest in intelligent monitoring, anomaly detection, and automation technologies
  • Excellent problem-solving skills and attention to detail.
  • Experience with API's and secure file transfer procedures.

Nice To Haves

  • Experience with telemetry standardization (OpenTelemetry) and observability data correlation.
  • Understanding of AI/ML concepts and their application to observability and operations (AIOps).
  • Experience with tools such as Ansible, Stonebranch, Kafka, and their role in system reliability.
  • Experience with CI/CD and developer platform tools such as Jenkins, GitHub, GitHub Actions, Artifactory, and Vault.
  • Familiarity with Snowflake and MongoDB Atlas, including experience with writing and executing basic queries, validating data, and troubleshooting production issues.
  • Understanding of Single Sign-On(SSO) technologies, including Microsoft Entra ID, SAML and OAuth.
  • Experience supporting enterprise applications in banking or financial services industry with understanding of regulatory, security and compliance requirements.

Responsibilities

  • Define and operationalize SLIs, SLOs, and error budgets
  • Own the incident management lifecycle (detection → triage → resolution → RCA → prevention)
  • Lead problem management and eliminate recurring issues
  • Develop and maintain runbooks, playbooks, and recovery procedures
  • Drive resilience engineering, including failover testing and capacity planning
  • Implement and optimize AIOps capabilities using platforms such as Moogsoft
  • Leverage Dynatrace for deep APM insights
  • Integrate alerting and escalation workflows with PagerDuty
  • Utilize synthetic monitoring via Catchpoint
  • Design and maintain centralized logging solutions using the ELK stack (Elasticsearch, Logstash, Kibana) and enterprise Logging as a Service (LaaS) platforms
  • Perform log analysis, correlation, and anomaly detection to support proactive issue identification
  • Drive event noise reduction and intelligent alerting strategies
  • Build dashboards and observability KPIs for operational insights
  • Design and implement automation-first solutions using Ansible and scripting (Python, Bash)
  • Enable self-healing capabilities (auto-remediation, restart logic, workflow recovery)
  • Orchestrate workflows using Stonebranch
  • Reduce operational toil through automation and continuous improvement
  • Provide L2/L3 support for data pipelines and integration workflows, including Snowflake ingestion and transformation processes
  • Support Snowflake pipelines, ETL workflows, and orchestration dependencies
  • Troubleshoot across distributed systems including: Object storage (e.g., S3), APIs, messaging, and file transfer systems
  • Support containerized workloads on OpenShift
  • Administer and support Windows Server and IIS-based applications
  • Manage certificate lifecycle (TLS, SAML, OAuth)
  • Support SSO integrations via Microsoft Entra ID
  • Work with relational databases (SQL Server, PostgreSQL) for troubleshooting and performance tuning
  • Lead and coordinate Disaster Recovery (DR) planning and execution
  • Validate failover processes across dependent systems
  • Ensure end-to-end DR readiness across integrated platforms
  • Document recovery strategies and participate in DR exercises
  • Ensure adherence to enterprise security, compliance, and audit requirements
  • Collaborate with IAM, PAM, logging, and cloud platform teams
  • Support change management, CAB processes, and release coordination
  • Provide reporting, KPIs, and executive summaries on system health

Benefits

  • bonuses
  • flexible benefits
  • competitive compensation
  • commissions
  • stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • A world-class training program in financial services
  • Opportunities to do challenging work
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service