Failure Analysis & System Validation Engineer

Advanced Micro Devices, IncAustin, TX
Hybrid

About The Position

As a Failure Analysis Engineer in Product Quality Organization, you will play a critical role in failure analysis across cutting-edge semiconductor products. This position is ideal for someone with a strong background in Post Silicon Validation, Failure Analysis, Platform level Debug who can apply deep technical insight to identify, isolate, and resolve complex issues in silicon and system-level environments. You will collaborate closely with cross-functional teams including design, validation, FW, and manufacturing to drive root cause analysis and corrective actions. Your contributions will directly impact product quality, reliability, and customer satisfaction.

Requirements

  • Strong technical and analytical expertise in Component Debug, Power Management, PCIe, and Memory Validation.
  • Ability to thrive in both collaborative team settings and independent work.
  • Ability to define goals, manage development efforts, and deliver high-quality solutions.
  • Bachelor's or master's degree in electrical engineering, Computer Engineering, or related field.

Nice To Haves

  • Experience in CPU/GPU Architecture, debug/validation, stress or functional test development.
  • Experience in Debugging complex Software/Hardware/FW related Fails from new product introduction (NPI) through production.
  • Experience working with High-speed server Interfaces such as PCIe, Memory and UPI. Ability to understand and debug PCIe and Memory related Failures.
  • Strong Programming experience with Python, shell script, or similar, Windows and Linux OS.
  • Hands on Experience with server platforms and lab equipment (Regulated Power supply, oscilloscopes, logic analyzers, etc.)
  • Understanding Firmware and drivers, and their interaction with Hardware. Ability to tune Firmware as needed.
  • Understanding of Power Management Concepts and Working knowledge of PVT dependencies and Margining (Voltage, Temperature, Frequency).
  • Strong Leadership, communication, documentation and presentation skills.
  • Experience with GPU data center infrastructure is a plus
  • Experience with AI tools, AI agents, or building agent-based automation workflows is a plus.

Responsibilities

  • Bring up Server Platforms, Develop and execute DOE's to reproduce and isolate failures, Identify and help resolve quality issues.
  • Develop Automation and tools to run tests and analyze results/logs.
  • Perform debug and drive root cause analysis across silicon, system and Rack-level issues
  • Collaborate with Customer Quality Engineers (CQE), design and validation teams to identify root causes for hardware and firmware issues.
  • Partner with Product Line Quality (PLQ)/Test teams to implement corrective actions and improve test capability.
  • Own end to end FA, Document findings, Analyze and present Data and contribute to continuous improvement initiatives. Create Failure Analysis reports.
  • Participate in NPI of Future products, Oversee the set-up of new products and test stations for Failure Analysis operations.

Benefits

  • AMD benefits at a glance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service