About The Position

This role focuses on making Luma's multimodal models fast by profiling and optimizing GPU, CPU, and accelerator code. The goal is to ensure efficient training and scalable deployment without compromising quality. The position involves writing kernels and operations to maximize hardware utilization. This is a deep performance-focused role requiring expertise in fused kernels, tensor cores, Triton, CUDA, and distributed multi-node deployment. A strong background in GPU optimization and a deep understanding of transformer internals are essential. Proficiency in CUDA, Triton, and profilers is a must.

Requirements

  • Expert-level Triton/CUDA programming and GPU optimization.
  • Strong PyTorch skills, including kernel development and custom operations.
  • Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling).
  • Deep understanding of transformer architectures and attention mechanisms.

Nice To Haves

  • Experience with compilers and exporters (torch.compile, TensorRT, ONNX, XLA).
  • Experience optimizing inference workloads for latency and throughput.
  • Triton compiler and kernel fusion techniques.
  • Knowledge of warp-level intrinsics and advanced CUDA optimization.

Responsibilities

  • Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency.
  • Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.
  • Develop fused kernels and leverage tensor cores and modern hardware features across platforms.
  • Optimize model architectures and implementations for distributed multi-node production deployment.
  • Build performance monitoring and analysis tools and automation.
  • Research and implement cutting-edge optimization techniques for transformer models.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service