About The Position

We are sharing a specialised consulting opportunity for experienced GPU Programming Software Engineers with strong expertise in CUDA, WebGPU, GLSL, C++, GPU architecture, kernel and shader optimisation, and high-performance parallel computing to contribute to an advanced AI training and GPU-programming evaluation project. Selected professionals will design and implement GPU-focused technical tasks, optimise kernels and shaders, analyse GPU performance, and review AI-generated solutions for correctness and efficiency. The work is suited to engineers with practical GPU-programming experience across graphics, machine-learning acceleration, scientific computing, HPC, or comparable GPU-intensive domains.

Requirements

  • Advanced professional experience with GPU programming
  • Strong proficiency with CUDA, WebGPU, GLSL, or another GPU technology capable of targeting NVIDIA hardware
  • Strong C++ proficiency
  • Deep understanding of GPU architecture and parallel execution
  • Experience profiling and optimising GPU kernels or shaders
  • Strong performance-engineering and debugging skills
  • Background in graphics, machine-learning acceleration, scientific computing, HPC, or another GPU-intensive domain
  • Experience analysing memory-access patterns and GPU resource utilisation
  • Ability to reason about architecture and performance trade-offs
  • Strong technical problem-solving and communication skills

Nice To Haves

  • Experience creating or reviewing rigorous programming problems is advantageous
  • No prior AI-training or model-evaluation experience is required

Responsibilities

  • Design and implement GPU workloads using CUDA, WebGPU, GLSL, or comparable technologies
  • Develop and optimise CUDA kernels and shader-based workloads
  • Apply appropriate parallelisation strategies and GPU execution models
  • Improve memory-access patterns, throughput, latency, occupancy, and resource utilisation
  • Ensure performance improvements preserve correctness
  • Profile GPU applications to identify computational and memory bottlenecks
  • Analyse thread organisation, synchronisation, execution behaviour, and architecture-specific constraints
  • Compare alternative GPU implementations and optimisation approaches
  • Identify inefficient or incorrect GPU execution patterns
  • Document performance findings and technical trade-offs clearly
  • Develop host-side logic and integrations in C++
  • Manage CPU–GPU communication, data transfer, and execution workflows
  • Integrate GPU functionality into broader software systems
  • Structure GPU workloads within maintainable application code
  • Debug performance and correctness issues across host and device components
  • Design technically rigorous GPU-programming tasks for AI training
  • Create problems testing GPU architecture, optimisation, and performance reasoning
  • Define clear expected outcomes and evaluation criteria
  • Review AI-generated GPU solutions for correctness, efficiency, and scalability
  • Provide structured technical feedback supporting model improvement

Benefits

  • Fully remote
  • Output-based compensation structure with payment made per task that meets project specifications
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service