About The Position

We are looking for a Principal Software Engineer to advance the reliability and observability architecture of a large-scale, globally distributed AI platform. You will set technical direction and solve complex challenges across runtime, routing, capacity, infrastructure, and specialized hardware. This role offers the opportunity to work at the intersection of distributed systems, cloud infrastructure, and AI inference. You will learn how advanced AI models are deployed and operated globally, collaborate with experts across the stack, and build foundational capabilities that directly improve customer experience. This is an ideal role for a technical leader who enjoys solving ambiguous, cross-system problems, mentoring engineers, and shaping engineering practices across an organization. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day, we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Nice To Haves

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Experience with AI/ML serving platforms, high-performance computing, accelerator-based infrastructure or other compute-intensive distributed systems.
  • Experience with distributed tracing, capacity management, load balancing, admission control, retries, backpressure, and graceful degradation.
  • Proficiency in Azure Monitoring systems in the Azure ecosystem would be a plus.
  • Experience designing, building, and operating large-scale distributed systems, cloud services, or other complex production platforms.
  • Experience with reliability engineering, observability, service health, telemetry, and production incident response.
  • Demonstrated ability to lead complex technical initiatives and drive alignment across engineering teams and organizational boundaries.
  • Strong written and verbal communication skills, including the ability to explain technical strategy and architectural decisions to engineers and senior leaders.

Responsibilities

  • Define and drive the architecture and technical roadmap for platform reliability, observability, and operational health.
  • Design unified health and telemetry capabilities that connect customer impact with application, capacity, dependency, infrastructure, and deployment signals.
  • Develop platform safeguards for overload protection, capacity management, routing integrity, configuration consistency, and automated isolation and recovery.
  • Establish engineering practices for production validation, progressive delivery, regression detection, fault testing, and automated rollback.
  • Advance end-to-end request tracing and diagnostics across distributed services, including routing, retries, failover, and asynchronous operations.
  • Lead cross-team architecture efforts, mentor engineers, and turn production learnings into reusable platform capabilities and measurable reliability improvements.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service