Incident Analyst/ Change Manager - In Office

WiproPhiladelphia, PA
Onsite

About The Position

The Analyst III is a senior operational leader who acts as Incident Commander during major outages, leads problem management efforts, and reviews/approves complex changes for operational readiness. This role partners closely with SRE and engineering teams on reliability strategy.

Requirements

  • 3–5 years in SRE, Operations, Incident Management, DevOps, or related fields.
  • Expert knowledge of monitoring tools, logging systems, and incident response.
  • Strong troubleshooting skills: networking, Linux/Windows servers, cloud services.
  • Strong communication and leadership skills in high-pressure situations.
  • Network configuration

Nice To Haves

  • Hands-on experience with cloud platforms (AWS, Azure, GCP).
  • Automation experience (Python, Go, Bash, or PowerShell).
  • Familiarity with microservices, containers, and distributed architectures.

Responsibilities

  • Lead high severity P1/P0 incidents as Incident Commander.
  • Coordinate cross functional engineering teams in real time.
  • Drive rapid troubleshooting, impact assessment, and resolution decisions.
  • Ensure high-quality incident documentation, executive-ready summaries, and follow through.
  • Apply Technical knowledge of Application architecture flows in driving the incident towards mitigation.
  • Lead problem investigations for major or recurring incidents.
  • Perform deep root cause analysis with engineering teams.
  • Validate corrective actions and track long-term problem remediation.
  • Present problem findings and preventive strategies to leadership.
  • Review and approve high-risk or complex changes for operational readiness.
  • Participate in Change Advisory Board (CAB) when needed.
  • Validate rollback strategies and operational safety measures.
  • Lead change execution for major maintenance or reliability events.
  • Drive reliability initiatives to reduce MTTA, MTTR, and incident volume.
  • Mentor junior analysts on incident handling and operational maturity.
  • Partner with SRE teams to expand observability, automation, and resilience.
  • Partner with SRE/Engineering teams on service reliability initiatives.
  • Lead maintenance events, failover tests, and resilience validation exercises.
  • Review and enhance runbooks, automation workflows, and monitoring strategies.
  • Contribute to Automation & AI Ideas for improving efficiency & reduce MTTM /MTTR.
  • Perform deep post incident analysis to identify systemic issues.
  • Contribute to automation solutions and self healing systems.
  • Own service reliability dashboards and operational KPIs.
  • Provide technical leadership to Analyst I/II team members.
  • Serve as a point of escalation for complex incidents.
  • Lead operational readiness reviews and training sessions.

Benefits

  • full range of medical and dental benefits options
  • disability insurance
  • paid time off (inclusive of sick leave)
  • other paid and unpaid leave options
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service