Software Engineer II

MicrosoftMountain View, CA
$102,100 - $219,200

About The Position

The AI Infrastructure team is responsible for building and operating large-scale, highly reliable, and efficient GPU infrastructure that powers Microsoft’s AI ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Microsoft 365 Copilot, GitHub Copilot, Microsoft Copilot, and Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads across the company. As a Software Engineer on the AI infrastructure team, you will work on cutting edge infrastructure and tools to support large scale model deployments, pre-training, post-training and fine-tuning on latest generation of NVIDIA and AMD GPUs in Azure and Microsoft partner clouds on some of the world’s largest AI Supercomputers. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Nice To Haves

  • 2+ years designing, developing, and shipping high quality software.
  • 2+ years of experience with distributed systems and cloud-based infrastructure.
  • 1+ year of experience with DevOps practices (CI/CD, automated testing, deployment, etc.).
  • 2+ years of software development experience in C#, C++, Python, or similar languages.
  • 2+ years of experience with containerization tools (e.g., Docker, Kubernetes).
  • Knowledge and hands on experience with production ML systems, large-scale training infrastructure, NCCL, CUDA libraries and tools

Responsibilities

  • Design, develop, and maintain AI infrastructure services in Go, Rust, Python, C++, and C#, deployed on large-scale Kubernetes clusters to support inference, pre-training, and post-training workloads for state-of-the-art AI models.
  • Collaborate with engineers, researchers, and external partners to troubleshoot issues, improve reliability, and optimize the performance of large-scale AI training and inference systems.
  • Build and enhance distributed systems that deliver high reliability, low latency, operational efficiency, and strong security across Azure and partner cloud environments.
  • Develop automation and tooling to improve GPU capacity utilization, streamline fleet operations, and enable efficient scaling of AI infrastructure.
  • Provide operational support, technical leadership, and vision while contributing to the deployment, monitoring, and continuous improvement of engineering systems and practices.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service