We are seeking an experienced engineering manager to lead a team of Site Reliability Engineers responsible for the availability, resiliency, security, and operational excellence of large-scale ad serving systems across hybrid on-premises and Azure environments. In this role, you will guide the design and operation of observability, automation, incident response, scaling, and failover capabilities; drive continuous improvement through blameless postmortems and infrastructure-as-code practices; and partner with machine learning, platform, and engineering teams to improve developer experience and accelerate reliable research-to-production workflows. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. Starting January 26, 2026, Microsoft AI (MAI) employees who live within a 50- mile commute of a designated Microsoft office in the U.S. or 25-mile commute of a non-U.S., country-specific location are expected to work from the office at least four days per week. This expectation is subject to local law and may vary by jurisdiction.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior