About The Position

NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA, you will have a meaningful role in keeping our Digital Marketing Services reliable, fast, and efficient. You'll use innovative technology and work alongside skilled professionals who continuously explore new possibilities.

Requirements

  • MS or BS in Computer Science/Engineering or a related field, or equivalent experience.
  • 8+ years’ experience supporting technical operations in a live-site production environment with a real passion for CDN automation, tooling, and infrastructure supporting AI applications.
  • Strong knowledge of the Kubernetes Platform, deployments, and cloud-native automation.
  • Proven strengths in problem-solving and root-causing issues, while continuously seeking ways to drive optimization, efficiency, and the bottom line.
  • Advanced level experience with scripting and development in Python, fully automating operational steps with “one-click” rapid solutions.
  • Key participation in the incident management process for early recognition of all service-impacting issues, accurate triage, partner communication, impact containment, service restoration, and post-incident follow-up.
  • SRE On call experience is a must.

Nice To Haves

  • Strong Akamai CDN Support skills and deep understanding of edge computing/edge AI.
  • Solid experience with the AWS Cloud Platform and Kubernetes as a platform.
  • SRE on-call experience.
  • Hands-on experience deploying and scaling Generative AI/LLM applications, integrating vector databases, or managing GPU-accelerated infrastructure.
  • Excellent communication, presentation, and analytical skills; the ability to communicate sophisticated infrastructure and AI concepts clearly across different audiences and varying levels of the organization.

Responsibilities

  • Build and deploy large-scale dynamic URL redirects using Akamai Edge Redirector Cloudlets for promotional efforts and site migrations.
  • Configure Akamai Forward Rewrite Cloudlets to map inbound requests to SEO-friendly paths.
  • Provide on-call support for production-grade applications, responding to incidents, prioritizing issues, and driving resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.
  • Author, test, and activate shared and non-shared Cloudlet Policies via the Akamai Cloudlets Policy Manager.
  • Maintain custom match criteria — including Geo, Device Characteristics, RegEx, and Query Strings — to ensure efficient origin offload and intelligent content delivery.
  • Quickly identify and address user-reported problems throughout the Digital Marketing Organization ecosystem.
  • On-board new applications, AI/ML services, and model endpoints on AWS Infrastructure.
  • Implement monitors, alerts, and SOPs to ensure early detection and accurate response to service-impacting issues, including tracking model drift and inference latency.

Benefits

  • Highly competitive salaries
  • Comprehensive benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service