Senior Incident and Problem Manager

Aviso WealthToronto, ON
CA$100,000 - CA$110,000Hybrid

About The Position

We’re looking to fill an opening for a Senior Incident and Problem Manager to join our Technology Ops & Support Partners team. Reporting to the Director, Service Delivery & Ops Governance, this role is responsible for leading end-to-end IT incident, and problem management practices across the department. The Senior Incident and Problem Manager will ensure timely service restoration, effective major incident response, thorough root cause analysis, and the successful implementation of corrective and preventive actions. This role will also drive continuous improvement of ITIL-aligned processes, reporting, and operational governance. Working closely with other teams, the Senior Incident and Problem Manager plays a critical role in strengthening operational resilience, reducing service disruption, and improving accountability. The role supports production readiness for significant releases, oversees adherence to change governance requirements, manages the lifecycle of standard, normal, and emergency changes, and escalates risks, exceptions, and control gaps to maintain service stability and release quality.

Requirements

  • Bachelor’s degree in Information Technology, Computer Science, or equivalent work experience is required
  • 5+ years of experience in IT operations, service management, incident management, or problem management is required
  • Proven experience leading major incidents in a 24x7 or high-availability enterprise environment is required
  • Strong working knowledge of ITIL incident, problem, change, and service management practices is required
  • Experience with ITSM platforms such as ServiceNow, Remedy, Jira Service Management, or similar tools is required
  • Strong understanding of infrastructure, application, cloud, network, and cybersecurity environments is required
  • Demonstrated experience facilitating post-incident reviews, root cause analysis, corrective action tracking, and problem remediation is required
  • Excellent written and verbal communication skills, including the ability to communicate clearly with technical teams, vendors, business stakeholders, and senior leaders, are required
  • Strong crisis management, prioritization, facilitation, and decision-making skills are required
  • Ability to remain calm, organized, and outcomes-focused during high-pressure service disruption events is required
  • Fluent communication skills in English are required

Nice To Haves

  • ITIL Foundation or higher certification is preferred
  • Experience supporting production readiness, release readiness, operational acceptance, or service transition activities is considered an asset
  • Experience in financial services, regulated environments, or other high-availability industries is considered an asset
  • Experience producing incident/problem dashboards, operational reports, and service review materials is considered an asset
  • Exposure to Azure, AWS, hybrid infrastructure, and modern monitoring/automation tools is considered an asset
  • bilingual skills in French are an asset

Responsibilities

  • Lead the end-to-end lifecycle for major incidents, including coordination, escalation, communication, resolution, and closure
  • Facilitate major incident bridges with technical teams, vendors, and business stakeholders to restore service quickly and minimize business impact
  • Own post-incident reviews, root cause analysis, and corrective action tracking to reduce repeat incidents
  • Manage and mature the problem management process, including recurring issue identification, problem backlog governance, and known error tracking
  • Support production readiness assessments for material releases by reviewing operational risks, supportability, monitoring, rollback planning, communication readiness, and incident response preparedness
  • Partner with engineering and application teams to identify recurring operational gaps and help mature practices that improve reliability, supportability, and service resilience
  • Produce incident and problem management metrics, including MTTR, SLA performance, incident trends, recurrence rates, and action closure status
  • Partner with application, infrastructure, cloud, service desk, cybersecurity, vendor, and business teams to improve service reliability and operational accountability
  • Maintain and improve incident management procedures, escalation paths, communication templates, runbooks, and governance practices

Benefits

  • Excellent health, dental and insurance benefits
  • Generous vacation time
  • fitness benefit
  • parental leave top-up options
  • Matching contributions to our retirement program
  • learning & development
  • education assistance program
  • Regular social events
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service