Staff Deep Learning Compiler Engineer

quadricBurlingame, CA
$170,000 - $230,000Onsite

About The Position

Quadric is building the world’s first General-Purpose Neural Processing Unit (GPNPU) architecture, bringing high-performance AI, DSP, and ML compute to edge devices. As a Deep Learning Compiler Engineer, you will design and optimize the compiler stack (MLIR, TVM, LLVM) that bridges state-of-the-art neural network frameworks directly to our proprietary hardware architecture. You'll play a critical role in unlocking peak hardware performance, low latency, and memory efficiency for edge AI workloads.

Requirements

  • BS, MS, or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • 8+ years Hands-on experience developing deep learning compilers or compiler infrastructures (e.g., MLIR, Apache TVM, LLVM, XLA, TensorRT, or Glow).
  • Strong proficiency in C++ (14/17/20) and Python, with solid fundamentals in data structures, algorithms, and object-oriented design.
  • Familiarity with modern AI/ML frameworks (PyTorch, TensorFlow, ONNX) and deep learning operator representations.
  • Solid understanding of computer systems, memory hierarchies, parallel processing, and CPU/GPU/NPU instruction execution.

Nice To Haves

  • Experience with low-level kernel optimization, SIMD/vector programming, and memory allocation strategy development.
  • Knowledge of model quantization methodologies (INT8, FP8, mixed-precision) and post-training/QAT optimization techniques.
  • Prior experience working on software stacks for custom AI accelerators, DSPs, or embedded architectures.
  • Experience with FPGA bring-up, hardware emulation, or cycle-accurate simulator development.

Responsibilities

  • Design, implement, and maintain compiler optimization passes targeting Quadric’s processor architecture using frameworks like MLIR, Apache TVM, or LLVM.
  • Develop lowering pathways from high-level machine learning frameworks (PyTorch, TensorFlow, ONNX) down to optimized low-level kernel code.
  • Optimize neural network performance for memory throughput, latency, compute unit utilization, and power consumption.
  • Implement graph-level optimizations, including operator fusion, layout transformation, quantization (INT8/FP16), and memory allocation strategies.
  • Analyze novel deep learning model topologies (Transformers, CNNs, Vision-Language Models) and extend compiler support for new operators and primitives.
  • Benchmark and profile end-to-end model performance to identify and resolve compiler bottlenecks.
  • Collaborate closely with hardware and micro-architecture teams to define instruction set extensions, hardware acceleration features, and compiler requirements.
  • Develop software simulators, functional models, and test benches to validate compiler correctness and generated binary performance.
  • Participate in hardware bring-up and validation efforts on FPGA and ASIC platforms.

Benefits

  • Medical, dental, and vision insurance from day one - Premiums covered at 99% for Employees
  • Company-paid life Insurance
  • Voluntary supplemental life insurance
  • STD + LTD insurance
  • Commuter support including parking or Caltrain reimbursement.
  • FSA + HSA
  • Equity with the business
  • Paid Parental Leave
  • 401(k) Retirement Plan
  • Flexible PTO
  • Winter holiday shutdown
  • Catered lunch each day in our office
  • Collaborative, low-ego culture with significant ownership and impact
  • A work culture focused on innovative disruption
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service