About The Position

We are seeking a Sr. Inference Engineer specializing in GPU Kernel Optimization to join our LLM Inference Performance Analysis and Optimization team. This role is crucial for maximizing the performance of every LLM inference operation. Our team develops infrastructure for silicon-measured kernel benchmarking, tooling for model-level performance projection, and agentic optimization systems that enhance GPU kernels at the assembly layer. We collaborate closely with NVIDIA's compiler, kernel, hardware, and framework organizations to identify bottlenecks and achieve measurable performance improvements. If you are passionate about driving GPU performance at the forefront of LLM inference, we encourage you to apply.

Requirements

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems (e.g., code generation, automated optimization, multi-step reasoning workflows).
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling experience with CUPTI, NSYS, and NCU, with a proven track record of attributing bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM, and a clear understanding of how kernel selection impacts model-level throughput and latency.
  • Working knowledge of GPU kernel optimization (CUDA, CUTLASS, Triton, or equivalent) and the ability to read PTX or SASS output.

Nice To Haves

  • Deep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).
  • Track record of shipping agentic systems end-to-end (tool invention, multi-agent orchestration, silicon-verified validation) within a performance engineering or kernel optimization context.
  • Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).

Responsibilities

  • Drive GPU kernel microbenchmarking to measure competing kernel implementations with real-silicon fidelity across the full configuration space required for production LLM deployments.
  • Conduct end-to-end model performance analysis, linking performance data to model-level serving economics, identifying high-value optimization opportunities, and developing optimization policies for production inference deployments.
  • Apply AI-driven analysis for agentic kernel optimization to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements.
  • Collaborate closely with compiler, hardware, kernel, and framework teams to deliver upstream improvements and production-grade performance gains.

Benefits

  • Highly competitive salaries
  • Comprehensive benefits package
  • Equity
  • Benefits detailed at www.nvidiabenefits.com/
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service