Major Incident Manager

ZENITH INFOTEK LLCVancouver, WA
$110,000 - $120,000

About The Position

The Major Incident Manager is responsible for orchestrating and managing high-impact incidents to ensure rapid service restoration and effective communication. This role involves crisis command, stakeholder management, and post-incident analysis to improve future response efforts.

Requirements

  • Experience in crisis command and incident orchestration.
  • Ability to lead Major Incident Command Bridges/War Rooms.
  • Proficiency in coordinating diverse teams (engineering, infrastructure, application, vendors).
  • Skilled in stakeholder and executive communication, providing clear and concise updates.
  • Experience managing escalation paths.
  • Understanding of Problem Management and Root Cause Analysis (RCA) transition.
  • Experience facilitating Post-Incident Reviews (PIRs).
  • Knowledge of ITIL Change Management and Emergency Change Requests (ECRs).
  • Familiarity with tracking KPIs like MTTD, MTRS, and SLA compliance.
  • Ability to refine incident management playbooks and workflows.

Responsibilities

  • Instantly invoke the Major Incident Management process upon notification of a Priority 1 (P1) or high-impact Priority 2 (P2) event.
  • Chair and lead the Major Incident Command Bridge / War Room, coordinating internal engineering, infrastructure, application, and 3rd-party vendor teams.
  • Maintain absolute focus on service restoration and workarounds rather than immediate root cause analysis.
  • Broadcast timely, clear, and non-jargon status updates to executive leadership, business unit heads, and key stakeholders at fixed cadences (e.g., every 15–30 minutes).
  • Manage escalation paths to engage senior technical leads or external suppliers when resolution stalls.
  • Ensure seamless transition of the incident to the Problem Management team for Root Cause Analysis (RCA) once service is restored.
  • Facilitate Post-Incident Review (PIR) sessions to capture timelines, evaluate response effectiveness, and document lessons learned.
  • Authorize and log Emergency Change Requests (ECRs) required for immediate fixes in compliance with ITIL Change Management.
  • Track key performance indicators, including Mean Time to Detect (MTTD), Mean Time to Restore Service (MTRS), and SLA compliance.
  • Continuously refine Major Incident playbooks, escalation matrices, and response workflows.

Benefits

  • 401(k)
  • Dental insurance
  • Vision insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service