Operations Resilience Director

Compass DatacentersDallas, TX
Onsite

About The Position

The Operations Resilience Director at Compass owns the strategic direction and continuous improvement of both incident management and business continuity planning — assessing current-state process, tooling, and governance holistically across Operations, and driving a roadmap that scales as the business grows. This role manages a 24x7x365 team out of two locations, balancing hands-on operational leadership with the ability to influence process and stakeholders beyond direct reporting lines, and manages the full lifecycle of operational incidents and continuity events from detection through resolution and root cause analysis. This role drives continuity and rapid response between our on-site Operations teams, Strategic Partners and Customers - and maintains KPIs and performance tracking to enable effectiveness of the workflows and tools. A key focus is also deploying AI and automation to increase efficiency and eliminate errors. Success in this role requires a balance of strong leadership, critical thinking, and a forward-looking approach to technology — with the ability to perform in a fast-paced, high-accountability environment.

Requirements

  • Strong leadership
  • Critical thinking
  • Forward-looking approach to technology
  • Ability to perform in a fast-paced, high-accountability environment
  • Experience managing a 24x7x365 team
  • Experience with incident management
  • Experience with business continuity planning
  • Familiarity with industry frameworks (e.g., ITIL, NIST)
  • Experience with risk assessments
  • Experience with Business Impact Analysis
  • Experience with crisis communications planning
  • Ability to influence stakeholders outside direct reporting authority

Nice To Haves

  • Deploying AI and automation to increase efficiency and eliminate errors

Responsibilities

  • Oversee 24x7x365 team operations across two locations
  • Responsible for maintaining 24x7x365 staff levels, and ensuring 100% coverage
  • Provide operational and tactical leadership
  • Responsible for ensuring all staff are properly trained
  • Prioritize and manage deliverables for the team to execute
  • Motivate, coach, and lead the team to achieve department goals
  • Drive cross-functional collaboration to ensure program success and a positive customer experience
  • Steward incident process and activities for high impacting incidents, across the portfolio
  • Own the incident severity classification framework (Sev 1-3) and corresponding escalation and notification matrix, including bridge command responsibilities for high-impact events
  • Manage customer relationship and ensure the appropriate level of transparency is consistently provided to process stakeholders (internal or external)
  • Capture Customer feedback for improvement opportunities and ensure notification / SLA requirements are followed
  • Monitor activities to make sure deliverables are being met and processes are being followed
  • Enable the collection, tracking, and recording of data for incident response & reporting requirements
  • Conduct annual reviews & updates Program/policy documents, and subsequent training content
  • Maintain KPIs to identify opportunities to eliminate errors through drills & tabletop exercises
  • Responsible for creating and maintaining the Ops ticketing standards & guidelines
  • Support evidence collection and inquiries for Operational audits
  • Assess incident management process, tooling, and governance holistically across Operations; identify structural gaps versus industry frameworks (e.g., ITIL, NIST)
  • Build and maintain a strategic roadmap to senior leadership, including any business case for investment
  • Influence adoption of improved process and governance with stakeholders and teams outside direct reporting authority
  • Coordinate & contribute to risk assessments (threat and vulnerability analysis) to inform Business Impact Analysis and continuity strategy
  • Steward and maintain the Business Continuity Planning (BCP) policy for Operations, including Business Impact Analysis with defined Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO), and periodic plan updates
  • Shepherd recurring BCP tabletop exercises in coordination with Incident Mgmt drills, ensuring incident response and continuity planning are tested
  • Contribute to the crisis communications plan for incidents and continuity events, defining notification protocols and messaging for executives, customers, and (where applicable) external stakeholders
  • Map critical vendor and supply chain dependencies (power, cooling, connectivity, and other third parties) and incorporate their continuity risk into the Operations Resilience strategy
  • Embodies Compass’ four Core Convictions in daily leadership — Humility In, Pride Out, Actions & Words Are One, Continuous Improvement of People, Processes and Systems, and We Ask Why
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service