Help build and operate the trusted platforms that power Microsoft 365's most critical compliance, security, and governance services. As a member of our team, you will work at the intersection of large-scale cloud engineering, service reliability, and operational excellence, helping ensure that enterprise and government customers can rely on Microsoft services every day. You'll collaborate with engineers across Microsoft to solve complex technical challenges, improve resiliency, and drive innovation through automation, modern cloud technologies, and data-driven operations. As a Site Reliability Engineer, you will help design, operate, and continuously improve large-scale Microsoft 365 and Purview services that support millions of users worldwide. You will partner with software engineers, service owners, and reliability teams to monitor service health, automate operational processes, investigate production issues, and implement engineering solutions that improve availability, performance, security, and customer experience. This opportunity will allow you to accelerate your cloud engineering expertise, develop deep knowledge of distributed systems and large-scale service operations, and build advanced skills in automation, observability, incident management, and AI-powered operational tooling. You will gain end-to-end ownership experience across the service lifecycle while helping shape the future of reliability engineering through automation, data-driven operations, and AI-enabled solutions. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level