Product Test Engineer - Machine Learning Hardware, RRL Technical Engineering

Amazon•Florence, KY

About The Position

Shape the future of AI infrastructure! Join RRL and develop system-level test solutions for our innovative ML acceleration hardware, deployed across our global server fleet. We're seeking highly skilled and motivated Product Test Engineers to join our RRL test & repair operation. In this role, you'll be at the forefront of validating and ensuring that test infrastructure is deployed and operating at required scale. You'll be responsible for designing and implementing comprehensive system-level test strategies that cover the full spectrum of our ML acceleration products, from individual components to fully integrated systems. This position requires a unique blend of hardware knowledge, software expertise, and systems thinking, as you'll be working at the intersection of custom silicon, complex firmware, and high-performance ML workloads. You'll collaborate closely with cross-functional teams including hardware designers, software engineers, and operations specialists to develop robust test solutions that can scale to meet the demands of RRL's global infrastructure. Your work will be crucial in identifying and resolving integration issues, optimizing system performance, and ultimately ensuring that our ML acceleration products meet the highest standards of reliability and efficiency in real-world data center environments. If you're passionate about pushing the boundaries of ML hardware testing and have a knack for solving complex system-level challenges, we want you on our team.

Requirements

2+ years of Linux systems administration and/or development experience
1+ years of work in at least two of these languages: Python, Java, Perl, PHP, Ruby or Bash/Shell experience
Bachelor's degree in Computer Science or other technical degree or related experience
Knowledge of networking fundamentals
Experience working in a 24/7 production environment
Experience in Linux systems administration and/or development
Experience working in at least two of these languages: Python, Java, Perl, PHP, Ruby or Bash/Shell

Nice To Haves

2+ years of site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration experience
1+ years of building scripts, tooling, and automation for large-scale computing environments experience
Knowledge of configuration management systems, such as Puppet, Chef, Ansible, or related systems
Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration
Experience in network capture and systems troubleshooting
Experience building scripts, tooling, and automation for large-scale computing environments

Responsibilities

Design and implement system-level test strategies for ML acceleration products
Develop comprehensive functional and performance tests for complete ML systems
Create and maintain scalable test infrastructure for high-volume product validation
Implement product bring-up and first-boot test procedures
Drive improvements in test coverage, product quality, and manufacturing efficiency
Collaborate with hardware and software teams to ensure end-to-end product validation
Analyze system-level test data to identify and resolve integration issues
Debug complex hardware/software interactions in a production environment
Develop and maintain documentation for system test procedures and manufacturing processes
Optimize test workflows to balance thoroughness with production efficiency

Benefits

health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage)
401(k) matching
paid time off
parental leave

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume