Senior/Staff AI Infrastructure Engineer, Inference & Optimization

DiDi LabsSan Jose, CA
$169,783 - $338,694

About The Position

We are seeking an experienced and mission-driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting-edge AI models across on-vehicle and cloud infrastructure. In this role, you will design high-efficiency inference pipelines, build system-level stability frameworks, and optimize hardware execution to ensure ultra-low latency and rock-solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real-time autonomous systems.

Requirements

  • Bachelor’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.
  • 3–8+ years of industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment.
  • Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools.
  • Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM).
  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.
  • Demonstrated ability to diagnose complex software-hardware integration issues and drive scalable, production-grade solutions.

Nice To Haves

  • Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.
  • Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high-throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.
  • Background in deep learning training frameworks (PyTorch) and practical experience with model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation.
  • Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices.

Responsibilities

  • Own the deployment, optimization, and resource scheduling of vehicle-side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.
  • Lead vehicle-side system stability initiatives, conducting independent root-cause analysis and driving resolution for complex, system-level performance bottlenecks and runtime anomalies.
  • Architect and scale service-oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.
  • Track and evaluate cutting-edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.
  • Establish system-level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.
  • Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Benefits

  • bonus
  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service