DL Performance Software Engineer - LLM Inference

NVIDIAToronto, ON
CA$135,000 - CA$220,000Hybrid

About The Position

We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference software, optimize GPU kernels, drive industry benchmarks, and work with state-of-the-art research techniques to improve serving efficiency. You’ll collaborate across inference performance, kernels, training, large-scale serving, and research teams to push the frontier of accelerated computing for AI.

Requirements

  • Bachelor’s, Master’s, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE).
  • 5+ years of industry experience in software engineering or equivalent research experience.
  • Strong programming skills in Python and one of C/C++, Go, or Rust.
  • Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, deep learning theories.
  • Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).
  • Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).
  • Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.

Nice To Haves

  • Experience developing major features and optimizations for LLM inference engines (e.g., vLLM, SGLang).
  • Hands-on work with LLM inference and training runtimes (deploying LLMs to production, large-scale LLM pre-training and RL), ML compilers and DSLs (e.g., Triton, CuTe, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).
  • Experience with speculative decoding training and runtime features: tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, EAGLE.
  • Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.

Responsibilities

  • Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features and serving runtime algorithms.
  • Profile and optimize the inference framework (vLLM) with methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.
  • Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.
  • Conduct and publish original research that advances the Pareto frontier in ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into production-grade, open-source software.
  • Develop, optimize, and benchmark GPU kernels (both hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service