We are seeking a Sr. Inference Engineer specializing in GPU Kernel Optimization to join our LLM Inference Performance Analysis and Optimization team. This role is crucial for maximizing the performance of every LLM inference operation. Our team develops infrastructure for silicon-measured kernel benchmarking, tooling for model-level performance projection, and agentic optimization systems that enhance GPU kernels at the assembly layer. We collaborate closely with NVIDIA's compiler, kernel, hardware, and framework organizations to identify bottlenecks and achieve measurable performance improvements. If you are passionate about driving GPU performance at the forefront of LLM inference, we encourage you to apply.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior