Software Engineer, GPU Performance

HeyGen•Los Angeles, CA

About The Position

HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We’re looking for a Software Engineer focused on GPU performance to make the systems behind these experiences faster and more efficient. You will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. This role is a fit for an engineer who enjoys understanding how software uses the hardware beneath it.

Requirements

  • Experience optimizing GPU-based AI workloads or high-performance computing systems.
  • Proficiency in Python and experience with PyTorch or a similar machine learning framework.
  • Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and data movement between CPU and GPU.
  • Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements.
  • Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly.

Nice To Haves

  • Experience with CUDA, Triton, or C++ GPU programming.
  • Experience optimizing video, image, audio, diffusion, or Transformer models.
  • Familiarity with multi-GPU inference, GPU interconnects, quantization, or large-scale model serving.
  • Experience building performance benchmarks or regression testing infrastructure.
  • Prior experience in a fast-paced technology environment.

Responsibilities

  • Use NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement.
  • Identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of each optimization.
  • Improve performance through batching, scheduling, memory management, and better GPU utilization.
  • Develop or integrate high-performance GPU kernels when existing implementations limit performance.
  • Build benchmarks and automated checks that catch performance regressions across representative video workloads.
  • Collaborate with AI researchers and infrastructure engineers to bring optimizations into production.
  • Measure the effect of changes on latency, throughput, cost, and output quality.

Benefits

  • Competitive salary and benefits package.
  • Dynamic and inclusive work environment.
  • Opportunities for professional growth and advancement.
  • Collaborative culture that values innovation and creativity.
  • Access to the latest technologies and tools.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service