Senior GPU Performance Software Engineer

IntelHillsboro, OR
Hybrid

About The Position

The Software and AI (SAI) organization is seeking a highly skilled software engineer to contribute to the development and low-level optimization of oneDNN, a complex, cross-platform, open-source performance library that serves as the foundation for deep learning applications. This is a low-level software engineering and hardware-acceleration role focused on developing highly optimized math primitives, parallel algorithms, and GPU kernels that power industry-leading AI frameworks on Intel hardware.

Requirements

  • BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field
  • 5+ years of professional software development experience with expert-level modern C++
  • 2+ years of hands-on experience in programming and kernel optimization on GPUs (via SYCL/DPC++, OpenCL, CUDA, or HIP), or at least 5+ years of similar low-level performance optimization experience on CPUs
  • Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (e.g., multi-threading, SIMD/vectorization)

Nice To Haves

  • Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels)
  • Hands-on experience with GPU assembly-level tuning or compiler optimization
  • Familiarity with parallel programming APIs such as OpenMP or oneTBB
  • Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks

Responsibilities

  • Develop high-performance GEMM, convolution, and attention kernels for AI workloads
  • Design scalable JIT and codegen infrastructure for GPU kernel generation
  • Implement fusion and memory-traffic optimizations to maximize hardware utilization
  • Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.)
  • Build analytical and empirical performance models for kernel dispatch and tuning
  • Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths
  • Co-design GPU primitives and kernel architectures for next-generation Intel GPUs
  • Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks
  • Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads

Benefits

  • Competitive pay
  • Stock bonuses
  • Health benefits
  • Retirement benefits
  • Vacation benefits
  • Stock programs
  • Quarterly bonuses
  • Highly flexible hybrid/remote working options
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service