AI Kernel Engineer

Majestic LabsLos Altos, CA

About The Position

Design and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution. Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA). Collaborate with compiler and runtime teams to integrate kernels into Triton and PyTorch pipelines. Profile and tune kernels using tools like Perfetto, VTune, Tracy, Nsight, or custom simulators. Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, MXFP/FP4, etc.). Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team.

Requirements

  • Degree at any level in Computer Science, Computer Engineering, or a related field from a recognized university.
  • Strong background in parallel programming (CUDA, Triton, RISC-V Vector (RVV), POSIX Threads, or OpenMP).
  • Deep understanding of memory layout, vectorization, thread/block scheduling, and cache behavior.
  • Experience programming wide-SIMD vector units and matrix/systolic engines.
  • Familiarity with the parallel-computing ecosystem and its high-performance libraries and primitives, such as BLAS/BLIS, dense linear algebra, and parallel scan/reduce/sort.
  • Skilled in performance analysis and parallel debugging using tools such as GNU Debugger and Nsight.
  • Hands-on experience profiling and optimizing compute or AI workloads (e.g., GEMM, softmax, attention).
  • Solid grasp of numerical stability, precision formats, and mixed precision arithmetic.
  • Proficiency in C++17 or higher, with strong knowledge of standard algorithms, data structures, and generic programming paradigms.
  • Collaborative work style with the ability to operate effectively in multicultural, cross-disciplinary environments.

Nice To Haves

  • Experience with optimization of irregular algorithms, such as graph computations or sparse numerical linear algebra, combining high-level data structure design with low-level SIMD and synchronization optimizations.
  • Familiarity with LLVM/MLIR.
  • RISC-V systems or bare-metal programming, including RISC-V vector and matrix extensions.

Responsibilities

  • Design and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution.
  • Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA).
  • Collaborate with compiler and runtime teams to integrate kernels into Triton and PyTorch pipelines.
  • Profile and tune kernels using tools like Perfetto, VTune, Tracy, Nsight, or custom simulators.
  • Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, MXFP/FP4, etc.).
  • Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team.

Benefits

  • We are stronger together, which is why we cultivate a safe, respectful, and collaborative environment where we follow the best ideas based on science and merit, not hierarchy.
  • Majestic is proudly an equal opportunity employer deeply committed to diversity of race, religion, gender, sexual orientation, background, and thought.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service