Principal Software Engineer - AI Powered Observability

MicrosoftRedmond, WA
$142,800 - $304,200Hybrid

About The Position

Join Engineering Operations (EngOps) – the organization driving operational excellence across the Microsoft Cloud to strengthen quality, reliability, security, and customer trust. As part of EngOps, you’ll design solutions that prevent issues before they happen, embed AI-powered automation, and turn signals into actions that deliver measurable customer impact. Our culture of empowerment, inclusion, and growth mindset defines how we work. Azure Reliability is driving transformation to AI-powered operations by building scalable ML infrastructure that enables autonomous, reliable, and secure cloud systems. We are looking for candidates that can combine deep technical expertise in MLOps with a proven ability to deliver measurable business impact through continuous learning, policy-driven governance, and responsible AI practices. Success in this role means advancing operational autonomy, quality, and security, while fostering collaboration and accountability across teams. Every day, customers stake their business and reputation on our cloud. You can help #EngOps keep them secure, resilient, and ready. We are looking to hire a Principal Software Engineer - AI Powered Observability, to take on this rare opportunity to shape autonomous outage detection, AI-driven reliability, and customer-impact protection across Microsoft at global scale. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. This role will require a minimum of three days in office.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Nice To Haves

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Awareness of, and ability to reason about, modern distributed software design patterns and cloud systems architecture, including microservices, containers, load-balancing, queuing, caching.
  • Experience with C#/Java/C/C++/Golang.
  • Experience in building, shipping and operating reliable solutions.
  • Experience with Artificial Intelligence (AI), Machine Learning (ML), risk assessment systems, or intelligent automation solutions.
  • Experience leading technical initiatives and influencing architecture, design, and engineering best practices across teams.
  • Demonstrated ability to drive operational excellence through automation, scalability, reliability, and continuous improvement.

Responsibilities

  • Partner across multiple product groups to apply subject-matter expertise in distributed systems design practices, interactions between cloud technology layers and components, basic dependencies at scale, and the code that defines infrastructures.
  • Lead by example and mentors' others to produce extensible and maintainable code used across products.
  • Develop and evangelize insights, best practices, and standards that can be applied to improve system, platform, and/or product development and operations across the business.
  • Drive continuous improvements in the architecture, code, features, operations and comprehensive use scenarios of products by leveraging end-to-end technical expertise.
  • Make improvements to the product fundamentals and architecture, share knowledge and code, always looking for ways to make what we build useful to multiple teams and products.
  • Demonstrates end-to-end expertise in distributed systems design, interactions between cloud technology layers.
  • Provide technical leadership in test maturity reviews, static analysis reviews, meetings, on-call rotations, and incident responses throughout product development and operations cycles.
  • Provides deep business and technical expertise as required to resolve major incidents.
  • Embody our Culture and Values

Benefits

  • Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service