Sr. Incident Management Engineer

MicrosoftRedmond, WA
$119,800 - $234,700

About The Position

Microsoft is seeking a Senior Incident Management (IcM) Engineer to lead the resolution of critical service and infrastructure incidents across Azure and cloud platforms. This role combines incident leadership, technical problem-solving, operational excellence, and reliability improvement. The ideal candidate has experience leading high-severity incidents, driving root cause analysis, influencing engineering teams, and improving the operational maturity of large-scale services. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Requirements

  • Master's Degree in Electrical Engineering, Computer Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience.
  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

Nice To Haves

  • Familiarity with HW (EE, thermal, mechanical) issues and high-level ability to identify and route to SME as appropriate
  • Experience in incident management, cloud operations, Site Reliability Engineering, infrastructure engineering, or related disciplines.
  • Experience leading complex technical incidents involving multiple teams.
  • Proven troubleshooting and analytical skills.
  • Proven communication and stakeholder management capabilities.
  • Experience with Microsoft IcM or other enterprise incident management systems.
  • Familiarity with KQL, Azure Monitor, Power BI, or operational analytics tools.
  • Experience automating operational processes using scripting or automation frameworks.

Responsibilities

  • Lead mitigation and resolution of high-severity service and infrastructure incidents.
  • Coordinate cross-functional engineering teams during live-site events.
  • Drive clear communication and decision-making during incident response.
  • Ensure timely escalation, mitigation, and customer impact reduction.
  • Lead post-incident reviews and root cause investigations.
  • Drive corrective and preventive actions to reduce recurring incidents.
  • Identify systemic reliability risks and recommend improvements.
  • Partner with engineering teams to improve service resilience.
  • Improve incident management processes, response playbooks, and escalation paths.
  • Track and analyze operational metrics and trends.
  • Drive improvements in incident response effectiveness and service availability.
  • Support operational readiness reviews and outage preparedness activities.
  • Influence engineering teams using data and operational insights.
  • Partner with service owners, SRE teams, datacenter operations, and engineering organizations.
  • Foster a culture of accountability, reliability, and customer focus.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service