Senior Software Engineering Manager

Microsoft•Redmond, WA
•Hybrid

About The Position

We are seeking an experienced engineering manager to lead a team of Site Reliability Engineers responsible for the availability, resiliency, security, and operational excellence of large-scale ad serving systems across hybrid on-premises and Azure environments. In this role, you will guide the design and operation of observability, automation, incident response, scaling, and failover capabilities; drive continuous improvement through blameless postmortems and infrastructure-as-code practices; and partner with machine learning, platform, and engineering teams to improve developer experience and accelerate reliable research-to-production workflows. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. Starting January 26, 2026, Microsoft AI (MAI) employees who live within a 50- mile commute of a designated Microsoft office in the U.S. or 25-mile commute of a non-U.S., country-specific location are expected to work from the office at least four days per week. This expectation is subject to local law and may vary by jurisdiction.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Nice To Haves

  • Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • 4+ years managing Linux and open source systems.
  • 4+ years people management experience.
  • 6+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.).
  • Experience migrating complex systems from on-prem to public cloud
  • Experience using Infrasturcture-as-code (e.g. Terraform, Bicep) to manage complex systems
  • Knowledge of CI/CD pipelines
  • Solid knowledge of distributed systems, networking, and storage.

Responsibilities

  • Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of hybrid on-prem and Azure ad serving systems
  • Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into platform performance
  • Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem environments.
  • Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements.
  • Ensure data privacy, compliance, and secure operations across serving environments.
  • Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service