Staff Engineer, ML Profiling Tools

Samsung Semiconductor•San Jose, CA
•Onsite

About The Position

The AGI (Artificial General Intelligence) Computing Lab is dedicated to solving the complex system-level challenges posed by the growing demands of future AI/ML workloads. Our team is committed to designing and developing scalable platforms that can effectively handle the computational and memory requirements of these workloads while minimizing energy consumption and maximizing performance. To achieve this goal, we collaborate closely with both hardware and software engineers to identify and address the unique challenges posed by AI/ML workloads and to explore new computing abstractions that can provide a better balance between the hardware and software components of our systems. Additionally, we continuously conduct research and development in emerging technologies and trends across memory, computing, interconnect, and AI/ML, ensuring that our platforms are always equipped to handle the most demanding workloads of the future. By working together as a dedicated and passionate team, we aim to revolutionize the way AI/ML applications are deployed and executed, ultimately contributing to the advancement of AGI in an affordable and sustainable manner. Join us in our passion to shape the future of computing!

Requirements

  • Bachelor's with 10+ years, or Master's with 8+ years, or PhD's with 5+ years of industry experience.
  • Abundant experience designing and implementing rich, interactive data visualizations and user interfaces using modern web technologies.
  • Strong background or interest in creating developer tools, IDE extensions, or complex diagnostic dashboards.
  • Proven ability to translate massive, unstructured, or multidimensional telemetry/profiling data into intuitive, human-readable visual representations.
  • Strong track record of designing workflows specifically for technical users (software engineers, data scientists, or researchers), focusing on minimizing cognitive load and streamlining root-cause diagnosis.
  • Hands-on experience with, or a strong curiosity about, systems profiling and tracing tools (e.g., Perfetto, Chrome Tracing, NVIDIA Nsight Systems/Compute, PyTorch Profiler, TensorBoard, eBPF).
  • Baseline understanding of computing architectures (CPUs, GPUs, TPUs, interconnects/networking) and typical performance bottlenecks (memory bandwidth, compute utilization, synchronization stalls).
  • Excellent problem-solving skills and ability to think critically and creatively.
  • You’re inclusive, adapting your style to the situation and diverse global norms of our people.
  • An avid learner, you approach challenges with curiosity and resilience, seeking data to help build understanding.
  • You’re collaborative, building relationships, humbly offering support and openly welcoming approaches.
  • Innovative and creative, you proactively explore new ideas and adapt quickly to change.

Nice To Haves

  • Working knowledge of modern machine learning frameworks (e.g., PyTorch, JAX, TensorFlow) and execution paradigms (LLM inference, pipeline/tensor parallelism, kernel dispatch).

Responsibilities

  • Develop and maintain highly interactive, large-scale and complex web-based visual analytics tools on AI/ML workloads.
  • Build interactive data visualization (flame graphs, call trees, timeline charts) capable of fluidly rendering massive, multi-dimensional infrastructure telemetry datasets.
  • Partner closely with compiler engineers and hardware architects to translate highly complex performance profiling metrics into clear, intuitive developer workflows.
  • Analyze and optimize the developer experience for machine learning performance profiling, designing intuitive tooling and workflows to help engineers diagnose and eliminate inference bottlenecks.
  • Communicate effectively with stakeholders, including users, partners, and management, to ensure that the systems are delivered on time and within budget.
  • Complete other responsibilities as assigned.

Benefits

  • Medical/Dental/Vision/401k
  • Charitable giving match
  • 4+ weeks of paid time off a year, plus holidays and sick leave
  • Stipend for fertility care or adoption
  • Medical travel support
  • Virtual vet care for your fur babies
  • On-demand apps and free confidential therapy sessions
  • Onsite Café and gym, plus virtual classes
  • Flexible environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service