Senior Staff Performance Engineer

Samsung Semiconductor•San Jose, CA
•Onsite

About The Position

The AGI (Artificial General Intelligence) Computing Lab is dedicated to solving the complex system-level challenges posed by the growing demands of future AI/ML workloads. Our team is committed to designing and developing scalable platforms that can effectively handle the computational and memory requirements of these workloads while minimizing energy consumption and maximizing performance. To achieve this goal, we collaborate closely with both hardware and software engineers to identify and address the unique challenges posed by AI/ML workloads and to explore new computing abstractions that can provide a better balance between the hardware and software components of our systems. Additionally, we continuously conduct research and development in emerging technologies and trends across memory, computing, interconnect, and AI/ML, ensuring that our platforms are always equipped to handle the most demanding workloads of the future. By working together as a dedicated and passionate team, we aim to revolutionize the way AI/ML applications are deployed and executed, ultimately contributing to the advancement of AGI in an affordable and sustainable manner. Join us in our passion to shape the future of computing! Location: Daily onsite presence at our San Jose, CA office / U.S. headquarters in alignment with our Flexible Work policy.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent industry experience.
  • Required industry experience: 10+ years with a BS, 8+ years with an MS, or 5+ years with a PhD in performance engineering, AI systems, distributed systems, high-performance computing, or a related area.
  • Ability to interpret workload traces, runtime telemetry, and performance data to identify bottlenecks and explain their underlying causes, including how workload characteristics interact with specific operations and mechanisms within software and hardware layers.
  • Knowledge of the LLM software stack, including serving and scheduling, attention and KV-cache management, kernel launch and memory-transfer overhead, and collective communication, and how these interact with accelerator compute throughput, memory hierarchy, memory bandwidth, and interconnect topology.
  • Experience characterizing agentic workflows, long-context processing, MoE models, or disaggregated inference deployments.
  • Experience profiling and optimizing AI workloads on NVIDIA GPU platforms using Nsight Systems and Nsight Compute, with analysis informed by GPU execution, memory hierarchy, and communication mechanisms.
  • Experience analyzing multi-node AI deployments, including synchronization overhead, load imbalance, communication patterns, and scaling behavior.
  • Experience with AI frameworks or serving systems such as PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, or Megatron-LM.

Nice To Haves

  • Candidates with additional experience and a record of technical leadership across teams may be considered at the Senior Staff level.

Responsibilities

  • Build and operate AI environments that reflect production workloads, including agentic workflows, distributed inference, disaggregated serving architectures, and MoE deployments.
  • Collect workload traces, runtime telemetry, and performance data across the software stack from AI applications.
  • Characterize and compare workloads across environments and platforms, identifying compute, memory, communication, and scheduling bottlenecks from applications and frameworks to runtimes, hosts, and devices.
  • Communicate findings to hardware architects, systems engineers, and software researchers through reports, presentations, and architecture reviews.
  • Define performance evaluation methodologies and benchmarking standards for adoption across hardware and software teams, and set the technical direction for workload characterization.

Benefits

  • Medical/Dental/Vision/401k
  • Charitable giving match
  • 4+ weeks of paid time off a year, plus holidays and sick leave
  • Stipend for fertility care or adoption
  • Caregiver support
  • Medical travel support
  • On-demand apps and free confidential therapy sessions
  • Onsite Café and gym, plus virtual classes
  • Flexible environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service