Failure Analysis Manager

INSPYR Solutions•Georgetown, TX
•Onsite

About The Position

The Senior Manager, Failure Analysis Engineering serves as the technical authority for system-level failure analysis and product reliability across the full manufacturing lifecycle, from New Product Introduction (NPI) through High Volume Manufacturing (HVM). This role leads the identification of complex failure mechanisms, defines structured root cause methodologies, and drives cross-functional resolution to improve product quality, manufacturing yield, and long-term reliability. The position requires deep expertise in server hardware architectures, failure physics, and data-driven analysis, combined with the ability to influence engineering, quality, and manufacturing organizations. The scope includes end-to-end failure analysis ownership for NPI (DVT/PVT readiness), Production (L6, L10, system-level testing), and field/customer returns (RMA/DOA). The technical scope covers CPU, Memory, Storage, Power, Networking, and Thermal systems. Cross-functional engagement includes Test Engineering, Product Engineering, Quality (PQE/MQE), Supplier Engineering, and Manufacturing. Additional scope involves data-driven reliability and failure trend analysis, influencing product design, test coverage improvements, and manufacturing process improvements.

Requirements

  • Deep knowledge of server hardware architectures, including: CPU, Memory, Storage, Power, Networking.
  • Strong expertise in failure analysis methodologies and root cause analysis techniques.
  • Experience with system-level debugging, including: Electrical issues, Firmware issues, Hardware/software integration issues.
  • Strong statistical analysis and data interpretation skills, including: Manufacturing yield, Reliability, Failure trends.
  • Understanding of manufacturing test flows, including: L6, L10, System-level testing.
  • Experience with technical and data analysis tools, including: Oscilloscopes, Logic analyzers, Diagnostic tools, Python, SQL, Power BI or equivalent data visualization/analysis tools.
  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering or a related technical field.
  • 10+ years of experience in Failure analysis, Reliability engineering, System debugging, Hardware engineering, or Related technical disciplines.
  • Proven experience supporting NPI through HVM transitions in complex hardware systems.
  • Demonstrated track record of solving complex, cross-domain technical problems.
  • Strong ability to influence engineering, quality, manufacturing, and supplier organizations.

Nice To Haves

  • Experience with GPU systems and liquid cooling.
  • Experience in hyperscale manufacturing environments.
  • Automation experience.
  • Six Sigma certification.
  • Experience supporting hyperscale or data center server environments.
  • Knowledge of reliability modeling, including: Weibull analysis, MTBF, HALT, HASS.
  • Exposure to DFX methodologies, including: DFR – Design for Reliability, DFT – Design for Test, DFM – Design for Manufacturing.
  • Experience automating failure analysis workflows and data pipelines.
  • Experience working with suppliers and supporting component-level failure analysis.

Responsibilities

  • Lead complex failure analysis (FA) and root cause analysis (RCA) to identify system-level failure mechanisms across server platforms.
  • Define and standardize failure analysis methodologies, tools, processes, and best practices.
  • Drive reliability strategy and influence NPI readiness, including DFR, DFT, and test coverage.
  • Establish failure trend analysis across yield, escapes, and field returns to enable data-driven decision-making.
  • Serve as an escalation point for critical quality issues and lead cross-functional technical problem solving.
  • Drive corrective and preventive actions across design, test, and manufacturing to eliminate repeat failures.
  • Improve test effectiveness, reduce NTF (No Trouble Found) loops, and strengthen feedback loops into engineering.
  • Mentor engineers and elevate failure analysis capabilities across the organization.
  • Identify systemic failure drivers and develop technical strategies to improve product reliability and manufacturing performance.

Benefits

  • Direct Hire
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service