About The Position

NVIDIA is at the forefront of the AI revolution, and our research is shaping the future of large language models. We are looking for a Principal Scientist to set the technical direction for synthetic data generation across NVIDIA's frontier model efforts. You will define and build open-source libraries within the NVIDIA NeMo ecosystem that generate synthetic datasets across text, code, structured, and multimodal data, feeding the pre- and post-training of LLMs such as Nemotron. This role combines hands-on software engineering with applied research in generative methods, and you will collaborate with research, engineering, product, and model teams as well as external labs.

Requirements

  • PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.
  • 15+ years of engineering and research experience in synthetic data generation, generative modeling, multimodal machine learning, or related areas.
  • Deep technical understanding of LLMs, how data shapes their pre-training, post-training, and RL stages, and inference frameworks such as vLLM or TGI.
  • Proven track record of developing or maintaining software libraries used by a broad developer community.
  • Experience building and optimizing scalable data pipelines for large-scale model training — throughput, distributed inference, and cost at cluster scale.
  • Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL or similar.

Nice To Haves

  • Significant open-source contributions in ML or data tooling, with community adoption.
  • Experience with multimodal generation or understanding (vision-language, document AI, video, or audio).
  • Experience generating data for agentic, tool-use, or reinforcement-learning post-training, including RL environment design.
  • Background in differential privacy, de-identification, or synthetic data for regulated industries such as healthcare, finance, or government.
  • Experience influencing model training decisions at frontier scale, or partnering directly with pre-training and post-training teams.

Responsibilities

  • Build and scale data generation pipelines using LLM-based methods combined with automated quality evaluation, resulting in datasets to improve both initial training and fine-tuning of LLMs such as Nemotron. These data pipelines cover reasoning, coding, structured output, and multimodal understanding.
  • Pioneer data generation for agentic and tool-use training: synthetic trajectories, multi-turn interactions, function calling, and executable environments for reinforcement learning, including reward modeling and verifiable-reward data.
  • Advance multimodal synthetic data generation — image, document, video, and audio — in partnership with NVIDIA's model teams.
  • Advance privacy-preserving and safe synthesis — differential privacy, anonymization, and de-identification — enabling model training on sensitive data in regulated domains.
  • Develop and maintain open-source libraries and SDKs with clean APIs and strong documentation.
  • Drive software excellence with modern tooling, architecture based on configuration, and professional Git/CI-CD.
  • Publish original research at top machine learning and AI conferences to maintain NVIDIA's technical leadership.
  • Mentor scientists and engineers across the team, raising the technical bar and growing the next generation of researchers.

Benefits

  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service