Join the EC2 Nitro Machine Learning Systems team to revolutionize accelerated computing in the cloud. We're seeking an exceptional Software Development Engineer to build and optimize the performance measurement infrastructure for some of the most computationally intensive AI/ML workloads on AWS. In this role, you'll establish EC2 as the definitive source for best-known-configurations across diverse ML applications including LLMs, multimodal models, and video generation workloads. Your expertise will directly influence future platform designs by translating performance insights from state of the art research and customer workloads into technical requirements for upcoming accelerated platform launches. Your impact will extend from low-level systems (CUDA, EFA, firmware) through ML frameworks to serving layers, requiring deep technical knowledge and the ability to communicate complex performance data as actionable business insights. This position offers the unique opportunity to shape the future of machine learning infrastructure at cloud scale while working at the intersection of high-performance computing, distributed systems, and machine learning technologies. EC2 Nitro Machine Learning Systems is responsible for development, operations, and maintenance of ML platforms for training and inference. We build and optimize infrastructure that powers some of the most computationally intensive AI/ML workloads. Our team creates reliable, high-performance systems that enable customers to push the boundaries of what's possible with ML. Working with us means having the opportunity to influence the future of supercomputing in the cloud while solving complex technical challenges at massive scale. We collaborate closely with customers and internal teams to continuously improve our platforms and deliver innovations that accelerate machine learning workflows.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level