Software Engineer II - AI Infrastructure

Microsoft•Redmond, WA
•$102,100 - $219,200

About The Position

The future of AI and cloud computing depends on highly reliable, scalable, and efficient distributed systems, and our team is building the platform that powers that future. As part of the AI Infra team, you will work on large-scale infrastructure that supports some of the world's most demanding cloud services and AI workloads. The infrastructure you will work on powers OpenAI and OSS model hosting, large-scale inferencing, and the backend for Microsoft Copilots — across the largest capacity fleets in the industry. You will help design and build foundational systems that enable reliability, performance, and operational excellence at global scale. As a Software Engineer II, you will design, develop, and operate distributed systems that power mission-critical cloud services. You will collaborate across engineering disciplines to build resilient, scalable, and highly available platforms, leveraging strong software engineering fundamentals and data-driven decision making. You will contribute throughout the software development lifecycle, from architecture and implementation to deployment, monitoring, and continuous improvement. This opportunity will allow you to: Accelerate your technical and career growth by solving complex distributed systems challenges at cloud scale. Develop deep expertise in large-scale infrastructure, reliability engineering, system architecture, and operational excellence. Hone your collaboration and leadership skills while working across teams to deliver high-impact solutions that serve millions of users and workloads.

Requirements

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • OOP (Object Oriented Programming) proficiency and practical familiarity with common code design patterns
  • 2+ years of experience with service development in a distributed environment, in a dev-ops role, including concurrency management and stateful resource management
  • Hands-on experience with public cloud services at the IaaS level

Nice To Haves

  • Master's Degree in Computer Science or related technical field AND 3+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 5+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Responsibilities

  • Design, develop, test, deploy, and operate large-scale distributed systems and platform services that deliver reliable, scalable, secure, and high-performance experiences.
  • Build and enhance infrastructure capabilities, resource management services, and workload orchestration solutions that optimize system efficiency, utilization, availability, and performance.
  • Develop software and platform features that support capacity management, service scalability, policy-driven decision making, workload placement, prioritization, and operational flexibility across distributed environments.
  • Collaborate with engineers, product stakeholders, and cross-functional partners to define technical requirements, influence architecture decisions, and deliver high-quality solutions that address customer and business needs.
  • Develop, test, and maintain control plane services written in C#, hosted on Kubernetes (AKS) clusters.
  • Analyze complex production issues, identify root causes, and implement sustainable solutions that enhance scalability, maintainability, efficiency, and service health. Provide operational support and DRI (on-call) responsibilities for the service.
  • Contribute to engineering best practices through technical design reviews, code quality initiatives, knowledge sharing, mentorship, and a culture of innovation, accountability, inclusion, and continuous learning.

Benefits

  • health_insurance
  • dental_insurance
  • vision_insurance
  • 401k
  • paid_holidays
  • flexible_scheduling
  • professional_development
  • learning_development_program
  • tuition_reimbursement
  • employee_discount_programs
  • disability_insurance
  • life_insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service