About The Position

The Lead Infrastructure Engineer is responsible for the operational health, reliability, and continuous improvement of enterprise messaging and collaboration platforms. This role provides technical and operational leadership for Microsoft-based communication and productivity services, ensuring stable, secure, and highly available experiences for employees across the organization. The successful candidate will lead major incident response activities, drive problem management initiatives, improve service reliability through automation and monitoring, and partner closely with Engineering, Product, Security, Infrastructure, and Support teams. This individual serves as a technical leader, operational strategist, and trusted advisor focused on operational excellence and service resiliency.

Requirements

  • 5+ years of Technology Infrastructure Engineering and Solutions experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 5+ years of Engineering experience in Messaging and Collaboration applications, including on-premise and M365 within a large enterprise
  • 5+ years of PowerShell scripting experience

Nice To Haves

  • MS-900 (M365 Fundamentals) or MS-700 (Teams Communications Administration)
  • AZ-900 (Azure Fundamentals)
  • Experience supporting hybrid messaging environments
  • Experience in regulated enterprise environments
  • Experience implementing Site Reliability Engineering principles

Responsibilities

  • Lead operational support for Microsoft 365, Exchange Online, Microsoft Teams, SharePoint Online, OneDrive, Microsoft Copilot, and related collaboration technologies
  • Ensure platform availability, performance, reliability, and operational readiness across enterprise environments.
  • Develop and maintain operational standards, runbooks, knowledge articles, and support procedures
  • Partner with engineering teams to transition new services, features, and platform changes into production support
  • Serve as an escalation point for high-severity production incidents and complex service degradations.
  • Lead Major Incident Management activities from detection through restoration and executive communication.
  • Coordinate technical response teams during service disruptions and drive clear ownership of next actions.
  • Conduct post-incident reviews and ensure corrective actions are documented, assigned, and tracked to closure.
  • Establish processes that reduce Mean Time to Detect and Mean Time to Restore.
  • Drive root cause analysis for recurring issues, chronic service degradations, and systemic operational risks.
  • Lead problem management activities to eliminate repeat incidents and improve platform resiliency.
  • Track operational trends and develop strategies to improve service health, supportability, and user experience.
  • Partner with engineering teams to improve platform architecture and operational readiness.
  • Design, optimize, and maintain monitoring strategies across critical collaboration services.
  • Develop actionable dashboards, alerts, trend reports, and service health views using tools such as Splunk.
  • Analyze logs and operational telemetry to identify service anomalies and emerging risks.
  • Reduce alert fatigue through tuning, suppression logic, event correlation, and automation.
  • Lead automation initiatives that improve operational efficiency, reduce manual effort, and strengthen auditability.
  • Develop scripts, workflows, and tooling to automate recurring operational tasks and validation activities.
  • Implement self-healing or guided remediation capabilities where appropriate.
  • Promote continuous improvement focused on reliability, scalability, efficiency, and customer experience.
  • Ensure compliance with IT Service Management processes for incidents, requests, problems, changes, and knowledge.
  • Manage operational work through ServiceNow and maintain high-quality documentation and ticket hygiene.
  • Support service-level objectives, service-level agreements, operational metrics, audits, and risk assessments.
  • Contribute to governance frameworks supporting enterprise collaboration services.
  • Provide clear, concise, and timely communications during operational events and service-impacting issues.
  • Create executive-level summaries of incidents, trends, risks, and service performance.
  • Partner with Product, Engineering, Security, Infrastructure, Support, and business stakeholders to establish priorities.
  • Build trusted relationships across technical and business teams.
  • Lead through influence, ownership, and technical expertise.
  • Mentor engineers and operations analysts through knowledge sharing and practical coaching.
  • Foster a culture of accountability, continuous learning, and customer-focused service delivery.
  • Drive collaboration across organizational boundaries and manage competing priorities during high-pressure events.
  • Act as a trusted advisor to leadership on service health, operational risk, and reliability strategy.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service