Senior AI Performance Architect

Microsoft•Boston, MA
•$119,800 - $261,000

About The Position

Do you want to be at the forefront of innovating the latest hardware designs to propel Microsoft’s cloud growth? Are you seeking a unique career opportunity that combines technical capabilities, cross-team collaboration, with business insight and strategy? Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we come together with a growth mindset, innovate to empower others, and collaborate to achieve our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. In alignment with our Microsoft values, we are committed to cultivating an inclusive work environment for all employees to positively impact our culture every day. Join the Systems Planning and Architecture (SPARC) team within Microsoft’s Azure Hardware Systems and Infrastructure (AHSI) organization, the team behind Microsoft’s expanding Cloud Infrastructure and for powering Microsoft’s “Intelligent Cloud” mission. Microsoft delivers more than 200 online services to more than one billion individuals worldwide, and AHSI is the team behind our expanding cloud infrastructure. We deliver the core infrastructure and foundational technologies for Microsoft's cloud businesses including Microsoft Azure, Bing, MSN, Office 365, OneDrive, Skype, Teams and Xbox Live. We are seeking a Senior AI Performance Architect to help shape the next generation of Microsoft AI systems. In this role, you will analyze OAI and MSI training and inference workloads, develop analytical models and tools to evaluate system performance, and validate key assumptions through targeted measurements on silicon. You will study interactions across accelerator architecture, memory, networking, and systems software; evaluate trade-offs in performance, scalability, utilization, and cost; and work with MSI and partner teams to guide hardware-software co-design decisions for cloud-scale AI systems.

Requirements

  • Master's Degree in Computer Science, Electrical Engineering, Computer Engineering, Applied Mathematics, or a related field AND 3+ years of technical engineering experience; OR Bachelor's Degree in a related field AND 5+ years of technical engineering experience; OR equivalent experience.
  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to, the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Nice To Haves

  • Doctorate in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 6+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience
  • Experience in AI systems performance, computer architecture, performance modeling, high-performance computing, distributed systems, or accelerator software.
  • Strong understanding of computer architecture and the performance interactions among compute, memory, communication, and software.
  • Experience with analytical performance modeling, simulation, profiling, workload characterization, or system performance analysis.
  • Hands-on programming experience in Python and at least one systems or accelerator programming environment such as C++, CUDA, Triton, ROCm/HIP, or an equivalent platform.
  • Experience analyzing AI training or inference workloads, including latency, throughput, capacity, utilization, and scaling behavior.
  • Experience with transformer-based models, large language model training and inference, attention, KV-cache behavior, mixture-of-experts, or speculative decoding.
  • Experience modeling or optimizing distributed execution strategies such as tensor, pipeline, expert, or data parallelism, or prefill/decode disaggregation.
  • Knowledge of accelerator architecture, HBM and memory hierarchies, scale-up and scale-out networks, PCIe or other high-speed I/O, kernels, compilers, and runtimes.
  • Experience correlating analytical or simulation models with silicon measurements and performing root-cause analysis of performance gaps.
  • Experience with data analysis and visualization for performance studies and architecture decisions.

Responsibilities

  • Analyze frontier training and inference workloads to identify key performance requirements and bottlenecks.
  • Develop and improve analytical models, simulators, and data-analysis tools for AI system performance.
  • Evaluate trade-offs in performance, scalability, utilization, capacity, and cost across accelerator, memory, network, and software configurations.
  • Study hardware-software interactions in large-scale AI systems and translate workload insights into architecture requirements.
  • Work with MSI and partner teams to evaluate hardware-software co-design options for training and inference.
  • Develop and run targeted microbenchmarks on silicon to measure compute, memory, communication, kernel, and synchronization performance.
  • Compare analytical model predictions with silicon and end-to-end workload measurements, identify gaps, and improve model accuracy.
  • Communicate findings and recommendations through clear data visualizations, technical reports, and architecture decision materials.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service