About The Position

NVIDIA is seeking a Senior Staff Platform Engineer to architect, build, and scale foundational infrastructure for some of our most demanding compute and AI/ML workloads. The role spans distributed systems, cloud, networking, content delivery, automation, and reliability engineering. Success means turning complex infrastructure challenges into resilient platforms that can grow with NVIDIA’s rapidly evolving needs. This is a hands-on technical leadership role with broad influence across Cloud, Networking, Security, AI/ML, and Developer Infrastructure teams.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience
  • 12+ years of relevant industry experience.
  • Proven success architecting, building, and operating large-scale distributed platforms or infrastructure systems in production.
  • Strong technical depth in several areas such as cloud infrastructure, distributed systems, networking, compute, storage, platform engineering, or content delivery.
  • Deep knowledge of Linux/Unix, TCP/IP, DNS, TLS, HTTP/S, proxies, load balancing, availability, scalability, and fault-tolerant system design.
  • Strong programming and automation skills with Python, Go, or similar languages, plus hands-on experience with infrastructure-as-code and orchestration.
  • Experience with AWS, Azure, or Google Cloud Platform and the ability to troubleshoot complex systems across application, operating system, network, and infrastructure layers.
  • Demonstrated ability to independently drive architecture and implementation across multiple teams, communicate effectively in complex situations, and mentor other engineers.

Nice To Haves

  • Experience building platforms for AI/ML training, inference, model serving, GPU-accelerated workloads, distributed compute, or high-performance computing.
  • Deep expertise with CDN and edge platforms such as Akamai, AWS CloudFront, Fastly, or Cloudflare, including caching, origin design, WAF, DNS, TLS, and global traffic management.
  • Experience developing self-service platform capabilities that enable engineering teams to consume infrastructure reliably and at scale.
  • Proven use of SLIs, SLOs, error budgets, capacity analytics, and reliability metrics to deliver measurable improvements.
  • Experience distributing models, datasets, containers, software artifacts, or other large objects across globally distributed environments.

Responsibilities

  • Architecting, building, and scaling highly available platform services for AI/ML, distributed compute, and data-intensive workloads.
  • Advancing CDN and edge infrastructure, including HTTP caching, origin architecture, TLS, WAF, rate limiting, traffic routing, and global load balancing.
  • Driving automation, infrastructure-as-code, and self-service capabilities using technologies such as Python, Go, Kubernetes, and Terraform.
  • Continuously improving reliability, scalability, efficiency, and operational simplicity using observability, capacity analytics, incident learnings, and performance data.
  • Collaborating with Cloud, Networking, Security, AI/ML, and infrastructure teams to solve complex problems that span multiple technology domains.

Benefits

  • Highly competitive salaries
  • Comprehensive benefits package
  • Equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service