AI Performance Modeling Engineer

quadric, IncBurlingame, CA
$150,000 - $200,000Onsite

About The Position

Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C++ code on a single programmable architecture. Our technology powers intelligent edge devices across automotive, industrial, robotics, and embedded systems. Founded by technologists from MIT and Carnegie Mellon, Quadric is a well-funded growth-stage semiconductor IP company with a growing licensing business. As we enter our next phase of growth, we're looking for our first true marketing leader to build and scale the function. Quadric has created an innovative General-Purpose Neural Processing Unit (GPNPU) architecture. Unlike standard accelerators, the Quadric GPNPU executes both neural network graph code and conventional C++ DSP/control code across edge and endpoint devices. As an AI Performance Modeling Engineer, you will build analytical, cycle-level performance models of AI inference workloads on our next-generation architecture in Python before silicon exists. These models directly guide team decisions on hardware lane bindings, tensor placement, and architecture trade-offs. We welcome candidates across all experience levels—from early-career engineers to seasoned experts—with direct mentorship provided to help you master mapping complex workloads (like LLMs) onto our custom hardware.

Requirements

  • Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code.
  • Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks (via industry experience, coursework, or research).
  • Comfort writing clear technical studies that state and defend evidence-based conclusions.
  • Deep understanding of NN inference operators and tensor shapes (e.g., Transformers, attention mechanisms, MoE, prefill/decode split) OR Proven performance modeling experience in another quantitative/technical domain.
  • BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.

Nice To Haves

  • Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels.
  • Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim).
  • Background in compiler internals (cost models, autotuners) or proficiency in C++.
  • Published performance studies or technical write-ups.

Responsibilities

  • Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware.
  • Derive from first principles which hardware lanes operations bind on (compute, on-chip/external memory bandwidth, interconnect) and model software pipelining overlaps.
  • Model tensor placement, tiling across processing elements, local memory residency, and data movement across memory tiers.
  • Model sharding and collective boundary communication across multi-die systems.
  • Incorporate architectural details across vision networks and Large Language Models (LLMs), including operator mix, sparsity, routing, and quantization/low-precision numeric formats.
  • Calibrate performance models against an instruction-set simulator and profiling traces to meet stated accuracy targets.
  • Write and defend technical studies presenting empirical evidence that directly informs architecture and product decisions.
  • Balance single-stream latency against scaled throughput performance.

Benefits

  • Medical, dental, and vision insurance from day one - Premiums covered at 99% for Employees
  • Company-paid life Insurance
  • Voluntary supplemental life insurance
  • STD + LTD insurance
  • Commuter support including parking or Caltrain reimbursement.
  • FSA + HSA
  • Equity with the business
  • Paid Parental Leave
  • 401(k) Retirement Plan
  • Flexible PTO
  • Winter holiday shutdown
  • Catered lunch each day in our office
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service