Senior Software Engineer - AI Infrastructure

Microsoft•Redmond, WA
•$119,800 - $234,700

About The Position

Help shape the platform that powers the next generation of artificial intelligence across Microsoft. As part of the Artificial Intelligence Infrastructure organization, you will build foundational platform services that enable the deployment, management, governance, and operation of large-scale Artificial Intelligence models and services across Microsoft. You will work at the intersection of cloud infrastructure, developer experience, reliability engineering, and Artificial Intelligence innovation, solving complex challenges that directly impact the speed, quality, and scale at which Microsoft delivers Artificial Intelligence products to customers worldwide. As a Senior Software Engineer, you will design, build, and operate highly scalable platform services that enable teams across Microsoft to onboard, deploy, monitor, govern, and manage Artificial Intelligence workloads throughout their lifecycle. You will collaborate closely with engineers, product leaders, researchers, and platform teams to deliver resilient cloud services, improve operational excellence, and accelerate the adoption of Artificial Intelligence infrastructure capabilities across Microsoft. This opportunity will allow you to: - Build foundational platform technologies that power some of Microsoft's most strategic Artificial Intelligence initiatives. - Develop deep expertise in large-scale distributed systems, cloud infrastructure, reliability engineering, and Machine Learning Operations platforms. - Expand your technical leadership skills by driving architecture, influencing engineering strategy, and partnering across multiple organizations to deliver high-impact outcomes.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Nice To Haves

  • Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Experience designing, building, and operating large-scale distributed systems or cloud services with high availability, reliability, and performance requirements
  • Experience with Microsoft Azure, Kubernetes, containers, infrastructure as code, and modern continuous integration and continuous deployment (CI/CD) systems.
  • Experience operating production services, including observability, incident management, capacity planning, and service health monitoring.
  • Experience building developer platforms, control planes, orchestration systems, workflow engines, or platform services used by other engineering teams

Responsibilities

  • Design, build, and operate large-scale distributed systems that power the lifecycle of Artificial Intelligence (AI) models and services across Microsoft.
  • Drive architecture, implementation, and evolution of cloud platform capabilities that improve deployment velocity, reliability, governance, and operational excellence for AI workloads.
  • Partner with engineering, product management, and research teams to translate complex business and technical requirements into scalable platform solutions.
  • Lead the development of highly available, secure, and observable services leveraging Site Reliability Engineering (SRE) principles and best practices.
  • Improve developer productivity by creating self-service experiences, automation, and tooling that simplify onboarding, deployment, monitoring, and management of AI applications.
  • Analyze production systems, identify performance and reliability opportunities, and drive continuous improvements through data-driven engineering decisions.
  • Mentor engineers, contribute to technical strategy, and influence engineering excellence through design reviews, knowledge sharing, and cross-team collaboration.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service