Manager Major Incident & Problem Management

Global Medical Response•Denver, CO
•Remote

About The Position

We are seeking a Manager to own the execution and continued maturity of two critical IT Service Management practices, Major Incident & Problem Management. This role leads the response to all enterprise Priority 1 incidents, coordinates rapid restoration of business services, and ensures leaders and stakeholders receive timely, accurate, and business-focused communications. Beyond active incident response, this position sets the strategic direction for Major Incident and Problem Management. The role provides functional, dotted-line leadership to the existing MSP (Managed Service Provider) Outage Coordinators and Problem Analyst, establishes consistent operating standards, drives accountability for root-cause and corrective-action work, and uses operational insights to reduce repeat incidents and improve service reliability. This is a hands-on leadership role for someone who can remain composed during high-impact events, bring structure to ambiguity, influence teams without relying on direct authority, and translate technical conditions into clear business impact and decisions.

Requirements

  • 5+ years of progressive experience in IT Service Management, service operations, incident management, problem management, or a related enterprise technology function.
  • Demonstrated experience leading high-severity incidents in a complex, multi-team environment with material business impact.
  • Experience designing, maturing, or governing Major Incident and Problem Management processes, not only executing individual cases.
  • Experience leading internal teams, managed-service providers, or matrixed resources through influence and clearly defined accountability.
  • Experience presenting incident status, risk, root cause, and corrective-action progress to senior technology and business leaders.
  • Strong working knowledge of ITIL practices, particularly Incident Management, Major Incident Management, Problem Management, Change Enablement, Configuration Management, Knowledge Management, and Service Level Management.
  • Practical experience with an enterprise ITSM platform; ServiceNow experience is strongly preferred.
  • Ability to understand complex application, infrastructure, network, cloud, integration, and vendor dependencies sufficiently to lead restoration and challenge assumptions.
  • Ability to use incident and problem data to identify trends, quantify operational risk, and prioritize improvement opportunities.
  • Comfort with on-call or after-hours engagement when enterprise P1 incidents require leadership.
  • Calm, decisive, and highly organized during fast-moving, high-pressure events.
  • Exceptional facilitation skills with the ability to maintain urgency without creating noise or confusion.
  • Clear writer and communicator who can translate technical detail into business impact, decisions, risks, and next steps.
  • Strong judgment, ownership, follow-through, and willingness to escalate when service restoration or corrective action is at risk.
  • Collaborative and credible with technical teams, business stakeholders, executives, and external partners.
  • Bachelor's degree in Information Technology, Computer Science, Business, or a related field, or equivalent practical experience.

Nice To Haves

  • ITIL 4 or 5 Foundation; ITIL Practice Manager, Monitor, Support and Fulfil, or equivalent advanced ITSM certification.
  • ServiceNow Certified System Administrator, Certified Implementation Specialist - IT Service Management, or equivalent platform experience.
  • Relevant incident command, problem analysis, reliability, or project leadership certification.

Responsibilities

  • Own and lead the end-to-end response for all enterprise P1 incidents, from declaration and bridge activation through service restoration, stakeholder transition, and formal closure.
  • Establish command and control during major incidents by clarifying roles, driving urgency, maintaining decision discipline, and ensuring the right technical and business resources are engaged.
  • Facilitate incident bridges, maintain focus on restoration, remove coordination obstacles, and escalate risks or resource gaps to technology leadership.
  • Ensure business impact, scope, workarounds, recovery progress, and restoration status are validated before they are communicated.
  • Coordinate executive, technology, and business communications using clear, concise, and audience-appropriate messaging.
  • Lead post-incident reviews and confirm that key decisions, timelines, lessons learned, and follow-up actions are documented.
  • Own the enterprise Problem Management practice, including intake, prioritization, investigation governance, known-error discipline, and closure criteria.
  • Ensure significant and recurring incidents are evaluated for problem records and that root-cause analysis is completed with appropriate rigor.
  • Drive accountable corrective and preventive actions with named owners, target dates, evidence of completion, and risk-based escalation for overdue work.
  • Partner with engineering, infrastructure, application, vendor, and service-owner teams to eliminate systemic causes and reduce recurrence.
  • Identify patterns across incidents, problems, changes, monitoring events, and service dependencies to inform reliability priorities.
  • Provide dotted-line leadership, operating direction, coaching, and quality oversight for MSP provided Outage Coordinators and the Problem Analyst.
  • Define role expectations, coverage models, escalation paths, facilitation standards, documentation requirements, and communication quality controls.
  • Conduct case reviews and targeted coaching to build consistency, confidence, and sound judgment across the team.
  • Coordinate workload and coverage with internal leaders and vendor management while maintaining clear accountability for practice outcomes.
  • Serve as the escalation point for complex incidents, stalled investigations, unresolved ownership, and process exceptions.
  • Develop and maintain the multi-year strategy, roadmap, operating model, policies, procedures, playbooks, and maturity plan for Major Incident and Problem Management.
  • Establish governance forums and performance reviews that focus on outcomes, risks, recurring failure themes, corrective-action health, and improvement priorities.
  • Define and monitor meaningful measures such as restoration performance, communication timeliness and quality, recurrence, root-cause completion, action aging, and business impact.
  • Identify opportunities to automate workflows, notifications, evidence capture, reporting, and handoffs across ITSM systems and adjacent platforms.
  • Align the practices with IT Service Management standards and integrate them with Change, Configuration, Knowledge, Event, Service Level, and Continuity Management.
  • Create training and simulation exercises that strengthen incident leadership, technical response, business-impact assessment, and executive communication.

Benefits

  • Comprehensive benefit options
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service