About The Position

NVIDIA’s accelerated computing platform is enabling the generational improvements in large language models, while the scale and complexity of these models are creating new challenges in computational efficiency. We are seeking a strong technical leader to drive a unified strategy for making LLMs more efficient from research through deployment. This role will bring together model innovation, systems expertise, and hardware awareness to ensure that new capabilities can be delivered within practical constraints of compute, memory, power, and cost. You will lead a multidisciplinary effort, establish the technical direction for LLM efficiency, and help shape how future models and computing platforms are designed together. The ideal candidate is a hands-on engineer who enjoys finding fundamental bottlenecks, challenging conventional boundaries between disciplines, and turning research ideas into scalable, real-world improvements.

Requirements

  • MS or PhD degree, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • 5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
  • Strong understanding of LLM architectures, training and inference workloads, and the tradeoffs between model quality, computational cost, memory footprint, latency, throughput, and power.
  • Strong background in performance analysis, roofline modeling, workload characterization, benchmarking, and hardware-aware optimization.
  • Proven ability to provide technical leadership and drive complex optimization projects from concept to production.

Nice To Haves

  • A track record of delivering measurable improvement throughput, cost per token, energy per token, memory efficiency, or time to train.
  • A first-principles - measure, model, optimize, and deliver - approach to improving LLM efficiency.
  • Familiarity with low-precision computation, quantization, sparsity, Mixture-of-Experts, long-context inference, and speculative decoding.
  • Experience co-designing model architectures with training, inference, compiler, or hardware constraints.
  • Experience influencing accelerator, system, or datacenter architecture based on future AI workload requirements.

Responsibilities

  • Lead cross-layer efforts to improve the efficiency of large language models across model architecture, training and inference systems.
  • Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure, and identify opportunities for model-system-hardware co-design.
  • Establish a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment.
  • Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence future model, software, and hardware roadmaps.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service