Principal Engineer - Major Incident Response & ITIL Platform

BJ's Wholesale ClubBJ's Club Support Center Marlborough, MA
Hybrid

About The Position

The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within ServiceNow and adjacent tooling platforms. Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.

Requirements

  • Bachelor's degree in Computer Science, Information Systems, or equivalent experience.
  • 8+ years in IT Service Management with a strong technical bias — hands-on platform work, not just process governance.
  • 5+ years of direct experience administering or engineering ServiceNow (workflow design, scripting, integrations, Performance Analytics).
  • Proven track record leading Major Incident Response in large-scale retail, e-commerce, or distributed digital environments.
  • Experience owning Problem Management and PIR programs with measurable outcomes (repeat incident reduction, MTTR improvement).
  • Demonstrated ability to manage and develop onshore/offshore teams in a follow-the-sun operations model.
  • Deep ServiceNow platform expertise understanding platform mechanics that drive configuration and administration.
  • Working knowledge of tools such as Dynatrace, Splunk, AlertOps/PagerDuty, or equivalent; ability to build event-to-incident automation bridges.
  • Observability & Monitoring: MS Teams, Jira — including integration design with ServiceNow.
  • Data & Reporting: Comfort with scripting (JavaScript, Python, or PowerShell) to accelerate toil elimination.
  • Able to command a major incident bridge — calm, decisive, technically credible under pressure.
  • Executive-level communication: concise, audience-aware, and trustworthy during crises.
  • Program management discipline: roadmaps, metrics, stakeholder alignment, and backlog ownership.
  • Growth mindset with a bias toward automation and measurable improvement.

Nice To Haves

  • ITIL 4 Strategic Leader or Managing Professional certification strongly preferred; ITIL Expert acceptable.
  • ServiceNow certifications (CSA, CIS-ITSM, or CIS-Event Management) highly desirable.
  • Collaboration & Comms: ServiceNow Performance Analytics, dashboard design, SLA/SLO instrumentation.

Responsibilities

  • Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
  • Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
  • Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
  • Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
  • Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within ServiceNow.
  • Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
  • Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within ServiceNow's CMDB and Problem Management modules.
  • Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
  • Configure and manage PIR workflows, SLA timers, and notification rules natively in ServiceNow — no manual handoffs.
  • Act as a hands-on technical owner of ServiceNow ITSM modules: Incident, Problem, Change, and Event Management.
  • Design and build ServiceNow workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, PagerDuty/AlertOps, etc.).
  • Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within ServiceNow Performance Analytics.
  • Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
  • Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
  • Ensure complete, accurate, and timely documentation of all Problems in ServiceNow with appropriate categorization and linkage to incidents and changes.
  • Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
  • Drive a culture of automation-first thinking: identify manual toil and eliminate it through ServiceNow scripting, Flow Designer, and third-party integrations.
  • Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
  • Present program health, metrics, and roadmap updates to senior IT and business leadership.

Benefits

  • Weekly Pay
  • Free BJ’s Memberships
  • Generous Paid Time Off
  • Flexible and Affordable Health Benefits
  • 401(k) Retirement Savings Plan
  • Employee Stock Purchase Plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service