About The Position

NVIDIA Cloud Functions team is looking for a motivated, product-minded AI/ML Engineer with domain expertise in AI platform engineering at scale. Our team builds and operates a serverless deployment platform for enabling AI applications. Our product enables and scales AI inferencing workloads using globally distributed orchestration of workloads on GPU-backed cloud-agnostic Kubernetes clusters. You will be working with a team of passionate and skilled engineers that are continuously innovating at the speed of light to provide the best product possible, for both external customers and internal NVIDIA teams. We are looking for someone to join us at the forefront of defining cloud engineering paradigms for AI at scale.

Requirements

  • Masters, PhD, or equivalent experience in Computer Science, Artificial Intelligence, Applied Math, or related field
  • At least 2 years work experience with Python, Rust, Golang, Linux or Bash.
  • Experience in Deep Learning and Machine Learning; expertise in using AI/DL frameworks and inferencing software such as SGLang, vLLM, TensorRT-LLM, or Dynamo.
  • Knowledge of CPU and GPU architecture.
  • Excellent interpersonal skills including ability to explain sophisticated technical topics to non-experts.
  • Experience in the design, implementation, and release of AI/ML products to market.
  • A flexible technologist familiar with all aspects of the software development lifecycle.

Nice To Haves

  • Demonstrate a strong desire to share knowledge with clients, partners and co-workers, able to show this through previous work.
  • Demonstrate expertise through projects or Open Source contributions in HPC, Data Analytics, Machine Learning, Deep Learning, Cloud Native Projects, Kubernetes, Slurm, or enabling GPU workloads.
  • Show a willingness and ability to dig into unfamiliar territories to tackle complex problems through examples in previous work.
  • Prior experience in building distributed systems.

Responsibilities

  • Becoming a trusted subject matter expert by understanding user challenges and constraints.
  • Translate this into product requirements and solutions, accelerating delivery of AI models and inference hosted on the NVCF platform.
  • Leading implementation of key features.
  • Conducting user-acceptance testing, load testing and performance evaluations.
  • Emphasis on customer experience, performance optimization and platform reliability.
  • Mentoring and embedding with other engineering teams building products on top of our platform on best practices for AI/ML workloads at scale with excellent performance, including ML reliability engineering at scale.
  • Producing reference architectures to guide customer use cases, applying the newest NVCF product features with the latest AI technologies.
  • Shepherding customer issues to resolution and providing timely warning of issues and risks.
  • Evaluating new and innovative technologies and tooling as the AI-at-scale landscape evolves to ensure we have a competitive product and forward-looking roadmap.

Benefits

  • highly competitive salaries
  • comprehensive benefits package
  • equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service