Principal Software Engineer

MicrosoftRedmond, WA
$142,800 - $304,200

About The Position

The FIT Infrastructure team builds the foundational accelerated compute platforms that power large‑scale AI training and inference across Azure. Our mission is to deliver secure, reliable, and highly efficient GPU infrastructure that enables multi‑tenant AI systems at global scale while maximizing utilization, performance, and developer productivity. This role sits at the intersection of cloud infrastructure, systems software, virtualization, and container platforms, working closely with Azure Infrastructure, OS, Networking, and Hardware teams to deliver end-to-end platform capabilities. You will work on mission critical infrastructure that directly powers largescale AI systems. Influence the future of cloud GPU platforms used by internal and external customers. Collaborate with experts across OS, hardware, networking, and AI platform teams. Opportunity to grow as a technical leader, shaping long term platform strategy.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Nice To Haves

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Proven ability to design and operate large‑scale, production infrastructure with high reliability and performance requirements.
  • Strong problem-solving skills and the ability to debug complex, cross layer systems issues.
  • Demonstrated technical leadership, including mentoring engineers and driving cross team architectural alignment.
  • Hands-on experience with virtualization and/or container platforms (e.g., VMs, Kubernetes, container runtimes).
  • Strong collaboration and communication skills, with the ability to work across organizational boundaries.
  • Experience in building or operating multitenant AI platforms in cloud environments.
  • Familiarity with high performance networking and low latency communication stacks.
  • Familiarity with GPU virtualization, passthrough, or partitioning technologies.

Responsibilities

  • Design and build GPU accelerated infrastructure for training and inference workloads, spanning bare metal, virtual machines, and containerized environments.
  • Develop systems for GPU device management, scheduling, isolation, and sharing (e.g., partial GPU allocation, multi‑tenant usage).
  • Build and operate advanced orchestration and resource governance scenarios using platforms such as AKS, Dynamic Resource Allocation (DRA), and related Kubernetes ecosystem capabilities to enable fair sharing, isolation, and efficient utilization of accelerated resources.
  • Build and evolve virtualization and container stacks to support modern AI workloads, including secure and confidential compute scenarios.
  • Optimize performance, reliability, and utilization across large GPU fleets, including scaleup and scale out configurations.
  • Partner with networking and storage teams to enable high performance interconnects (e.g., RDMA/InfiniBand class networking) for distributed workloads.
  • Drive end-to-end platform features from design through production, including observability, diagnostics, and operational excellence.
  • Influence platform architecture and technical direction across teams through design reviews and technical leadership.

Benefits

  • Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service