GPU AI Solution Architect

Intel Corporation•Hillsboro, OR
•$195,200 - $275,580•Hybrid

About The Position

As an AI Systems and Solutions Engineer, you will play a pivotal role in designing, optimizing, and supporting cutting-edge AI accelerator systems and cloud infrastructure to enable large-scale machine learning workloads. You will be responsible for system-level engineering, validation, and integration of next-generation AI platforms, ensuring optimal performance and reliability for enterprise applications. This role has a direct impact on advancing Intel's AI solutions and driving innovation in the rapidly evolving AI technology landscape. You will be joining Intel's Data Center Group (DCG), a global team dedicated to developing high-performance computing solutions for enterprise, cloud, and edge workloads. The AI Customer Success and Solutions organization focuses on delivering advanced technologies and systems that accelerate artificial intelligence adoption and maximize its potential across industries. The team collaborates broadly across Intel to deliver impactful solutions that empower businesses to solve complex challenges.

Requirements

  • Bachelors and 6+ years or Masters and 4+ years or PhD and 2+ years in Computer Science, Electrical Engineering, or related field.
  • 5+ years of experience in system engineering, platform validation, or related roles.
  • 3+ years of experience of successfully bringing up and debugging high-performance AI clusters.
  • 3+ years of experience resolving complex system-level issues in production AI/ML environments.
  • 3+ years of experience AI cluster design, validation, and production deployment experience.

Nice To Haves

  • Experience of debugging large GPU clusters is a plus.
  • Solid understanding of PCIe, memory subsystems, and AI accelerators.
  • Experience with Intel platforms (Xeon, Gaudi) or other GPU/AI accelerators.
  • Familiarity with AI frameworks such as PyTorch, TensorFlow, OpenMPI, and vLLM.
  • Proficiency in programming languages, especially Python.
  • Expertise in Linux/Unix administration, Docker, and shell scripting.
  • Knowledge of Redfish, IPMI, BMC management protocols.
  • Strong grasp of computer architecture, AI/ML workload optimization, and enterprise platform security.

Responsibilities

  • Design and optimize AI accelerator systems, including GPU clusters and Gaudi platforms, for production machine learning workloads.
  • Debug system-level issues such as PCIe connectivity, memory subsystems, and interconnects in AI clusters.
  • Lead platform bring-up and validation for next-generation AI hardware, ensuring system stability and readiness.
  • Develop and execute comprehensive test plans and validation strategies for AI systems.
  • Collaborate with OEM vendors on firmware integration and system-level optimizations.
  • Perform full-stack debugging across hardware, firmware, and software layers.
  • Develop automated testing frameworks, diagnostic tools, and monitoring solutions for AI systems.
  • Mentor junior engineers and contribute to cross-functional team collaborations to resolve complex technical challenges.
  • Drive architectural improvements and technical decisions to enhance AI infrastructure capabilities.

Benefits

  • competitive pay
  • stock bonuses
  • health
  • retirement
  • vacation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service