Senior AI Systems and Algorithms Engineer

NVIDIASanta Clara, CA
$152,000 - $287,500

About The Position

NVIDIA is seeking a Senior GenAI Algorithms Engineer to advance the state of the art in foundation model development, training, and deployment. You will work at the intersection of large-scale distributed training, reinforcement learning for LLMs/VLMs, model efficiency, multimodal AI, and open-source AI infrastructure. This role spans the entire GenAI lifecycle from large-scale data preparation to training, post-training, inference optimization, and framework development. You will collaborate with research, product, and infrastructure teams to design new algorithms, optimize existing systems, and contribute to NVIDIA's open-source AI stack, including Megatron-LM, Megatron Bridge, and NeMo-RL.

Requirements

  • MS or Ph.D in Computer Science, AI, Applied Mathematics, or a related field (or equivalent experience).
  • 5+ years of relevant industry experience.
  • Strong foundation in machine learning, deep learning, and optimization.
  • Excellent software engineering skills, including Python and PyTorch.
  • Experience building high-performance software for large-scale AI systems.
  • Strong analytical, debugging, and performance optimization skills.
  • Excellent communication and collaboration skills.

Nice To Haves

  • Distributed training at scale, including Megatron-LM, Megatron Bridge, FSDP, TP/PP/CP/DP, heterogeneous or per-module parallelism, optimizer research, and efficient sparse or long-context attention.
  • Supervised fine-tuning (SFT), reinforcement learning for LLMs (e.g., PPO, GRPO, asynchronous RL), and large-scale RL frameworks such as NeMo-RL.
  • Model compression techniques including quantization (FP8, NVFP4, INT4), pruning, knowledge distillation, neural architecture search, and diffusion or non-autoregressive language models.
  • Contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers along with experience in GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads.

Responsibilities

  • Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training.
  • Develop algorithms and systems that improve the scalability, efficiency, and cost of large-scale pre-training and post-training.
  • Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms.
  • Develop reusable infrastructure and contribute brand new model support to NVIDIA's open-source GenAI training platform.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service