About The Position

We are sharing a specialised consulting opportunity for experienced CUDA Engineering Experts with strong expertise in CUDA, C++, GPU kernel optimisation, GLSL, WebGPU, profiling, and high-performance computing to contribute to an advanced GPU-engineering and AI training project. Selected professionals will analyse and optimise GPU kernels, improve CUDA and C++ codebases, work with shader and WebGPU workflows, and document performance improvements across modern GPU architectures. The work requires strong low-level performance engineering, rigorous profiling, and the ability to translate technical findings into clear, actionable recommendations. No prior experience in AI is required.

Requirements

  • Demonstrated expertise in CUDA programming
  • Strong track record of GPU kernel performance optimisation
  • Advanced C++ development experience
  • Experience working in high-performance computing environments
  • Hands-on experience with GLSL and WebGPU
  • Strong understanding of graphics or compute shader development
  • Proficiency with GPU profiling tools such as Nsight, Visual Profiler, or equivalent
  • Ability to reason about memory behaviour, kernel execution, and hardware utilisation
  • Strong analytical skills for evaluating performance across GPU architectures
  • Experience refactoring performance-critical codebases
  • Excellent written and verbal technical communication skills
  • Comfortable collaborating in remote, cross-disciplinary environments

Nice To Haves

  • No prior AI-training experience is required

Responsibilities

  • Analyse and optimise GPU kernels using CUDA
  • Identify performance bottlenecks affecting computational throughput
  • Improve kernel efficiency across modern GPU hardware
  • Apply targeted optimisation techniques based on profiling results
  • Validate performance gains using quantitative benchmarks
  • Profile GPU workloads using tools such as Nsight, Visual Profiler, or comparable platforms
  • Diagnose memory, compute, occupancy, and execution bottlenecks
  • Evaluate kernel performance across different hardware generations
  • Develop data-driven optimisation strategies
  • Document measurable changes in latency, throughput, and resource utilisation
  • Refactor CUDA and C++ code for improved maintainability and efficiency
  • Improve architecture and organisation within performance-critical codebases
  • Reduce unnecessary complexity while preserving functionality
  • Adapt implementations for portability across GPU architectures
  • Apply high-performance computing best practices throughout development
  • Implement shader logic using GLSL
  • Develop graphics and compute workflows using WebGPU
  • Integrate shader-based processing into existing systems
  • Evaluate performance trade-offs across GPU execution environments
  • Maintain compatibility and consistency across pipeline components
  • Contribute expertise to GPU architecture and performance-design discussions
  • Evaluate new GPU-based implementation approaches
  • Assess technical trade-offs across performance, maintainability, and scalability
  • Support definition of meaningful performance metrics
  • Recommend practical approaches based on profiling and benchmarking evidence
  • Document optimisation strategies, benchmark results, and technical findings
  • Produce clear reports describing performance improvements
  • Communicate complex GPU behaviour to technical stakeholders
  • Collaborate with remote and cross-disciplinary project teams
  • Share relevant developments in GPU programming and performance engineering

Benefits

  • Independent contractor engagement
  • Fully remote
  • Compensation: $60–$100/hour
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service