Problem Manager (ITIL, ServiceNow, Root Cause Analysis)

Conduent Business Services, LLC, Remote US
$70,956 - $92,150

About The Position

Conduent is seeking an experienced Problem Manager to lead enterprise-wide Problem Management, Root Cause Analysis (RCA), and Incident Prevention initiatives across a complex technology environment. This role is ideal for a hands-on IT operations professional with expertise in ITIL Problem Management, Incident Management, ServiceNow, and service reliability improvement. You will work across infrastructure, cloud, network, application, and operations teams to identify recurring issues, drive permanent corrective actions, and reduce repeat incidents. Success in this role means delivering stronger RCAs, improving remediation execution, enhancing operational governance, reducing incident recurrence, and helping the organization build a proactive, data-driven Problem Management function.

Requirements

  • 7+ years of experience in IT Operations, Problem Management, Incident Management, Major Incident Management, or related disciplines.
  • Strong understanding of ITIL Problem Management, Incident Management, and operational governance.
  • Experience conducting Root Cause Analysis (RCA), 5 Whys investigations, and post-incident reviews.
  • Hands-on experience with ServiceNow or comparable IT Service Management (ITSM) platforms.
  • Strong analytical, documentation, facilitation, and stakeholder management skills.
  • Ability to influence cross-functional teams and drive accountability without direct authority.
  • Experience creating executive-level communications, dashboards, and operational reporting.
  • Infrastructure: Windows and Linux platforms, VMware and Hyper-V virtualization, Performance troubleshooting (CPU, memory, storage, and I/O)
  • Storage: SAN and NAS environments, Storage performance, redundancy, and resiliency concepts
  • Networking: Routing, switching, VLANs, DNS, Firewalls, load balancing, and packet flow analysis, Network latency and packet-loss interpretation
  • Cloud Technologies: AWS and Microsoft Azure, Cloud networking, compute, identity, and storage services, Cloud architecture and failure-mode analysis
  • Application & Platform Technologies: APIs, microservices, and distributed systems, CI/CD and deployment pipeline fundamentals, Understanding of how application defects, configuration drift, and dependencies create operational incidents

Nice To Haves

  • ITIL Foundation or higher certification.
  • Experience supporting large-scale enterprise environments.
  • Experience with operational analytics, trend analysis, and KPI reporting.
  • Exposure to automation, AI-enabled operational workflows, or AIOps practices.
  • Experience working with executive stakeholders and client-facing incident communications.

Responsibilities

  • Lead Root Cause Analysis (RCA)
  • Facilitate timely 5 Whys Root Cause Analysis (RCA) sessions for Priority 1 (P1) incidents and recurring Priority 2 (P2) incidents.
  • Review system logs, monitoring tools, dashboards, and change records to establish incident timelines and identify contributing factors.
  • Utilize AI-assisted tools to accelerate incident investigations, timeline generation, RCA documentation, and knowledge capture.
  • Document accurate root causes and corrective actions within ServiceNow Problem Management records.
  • Ensure RCA deliverables are factual, blameless, technically accurate, and suitable for leadership and client-facing communications.
  • Lead the end-to-end Problem Management lifecycle from investigation through closure.
  • Partner with technical teams to develop effective Remediation Action Plans (RAPs).
  • Ensure remediation efforts include both corrective and preventive actions with clear ownership and target dates.
  • Track remediation progress, escalate risks, and hold stakeholders accountable for commitments.
  • Validate that implemented solutions address root causes and prevent future recurrence.
  • Measure and report improvements in service reliability, operational risk, and incident recurrence.
  • Identify recurring patterns across infrastructure, cloud, application, and network incidents.
  • Drive proactive problem identification and long-term service improvements.
  • Recommend monitoring enhancements, automation opportunities, and architectural improvements.
  • Track key operational metrics including: Incident recurrence rate, MTTR (Mean Time to Resolution), Action-item closure rate, Chronic risk themes, Problem backlog health.
  • Facilitate weekly Problem Review meetings.
  • Maintain executive dashboards and Problem Management reporting.
  • Deliver concise, executive-ready summaries and status updates.
  • Produce high-quality RCA reports for customers and senior leadership.
  • Continuously improve governance processes, operational standards, runbooks, and reporting practices.
  • Champion automation and AI-enabled workflows that improve operational efficiency and insight generation.

Benefits

  • health insurance coverage
  • voluntary dental and vision programs
  • life and disability insurance
  • a retirement savings plan
  • paid holidays
  • paid time off (PTO) or vacation and/or sick time
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service