Staff Modeling Architect

NeurophosAustin, TX
Onsite

About The Position

Neurophos is developing a novel AI compute architecture using silicon photonics and active, programmable metasurfaces to perform matrix multiplications at the speed of light. This approach aims to overcome the energy limitations of traditional silicon by offering significantly higher energy efficiency and performance for large-scale AI inference. The company has secured $110M in Series A funding from prominent investors and is seeking a Staff Modeling Architect to join their world-class team. This role is crucial for bridging the gap between production models and performance/energy metrics, as well as creating functional models for pre-tape-out software development. The position involves close collaboration with architects, RTL and physical design teams, and compiler and runtime engineers, focusing on hardware/software co-design for emerging AI workloads.

Requirements

  • BS, MS, or PhD in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
  • 8+ years of experience in hardware modeling, functional modeling, performance modeling, performance simulation, or accelerator performance analysis.
  • Proven track record of shipping a model or study that was depended upon by other teams (architecture, compiler, customer, or silicon).
  • Judgment to select the appropriate method for a given question (roofline, limiter analysis, analytical models, trace-driven simulation, TLM, RTL simulation).
  • Strong understanding of computer architecture, microarchitecture, memory systems, and AI accelerators (GPU, TPU, NPU, custom SoC).
  • Proficiency in modern C++ (C++17 or later) for functional models, performance models, and simulation infrastructure.
  • Proficiency in Python for models, analysis, and plots, including NumPy, Pandas, and Matplotlib.
  • Experience working within a discrete-event, cycle-approximate, or cycle-accurate simulator (e.g., SystemC, gem5, SST, or custom kernel).
  • Ability to construct an LLM or accelerator workload from a model card or paper, covering prefill and decode, MoE, GEMM tiling, and quantization.

Nice To Haves

  • PhD in Computer Engineering, Electrical Engineering, or Computer Science.
  • Experience in hardware/software co-design, including work with MLIR, TVM, XLA, ONNX, operator fusion, or graph compilers.
  • Experience modifying or extending a simulation kernel.
  • Experience correlating an analytical model against silicon, vendor datasheets, or measured datacenter GPUs and inference accelerators.
  • Familiarity with TLM 2.x, Verilator, SystemVerilog, DPI, or UVM.
  • Familiarity with HBM, DRAM controllers, cache, SRAM, network-on-chip (NoC), AXI, DMA, and scratchpad memory.
  • Experience with power modeling using tools like McPAT, CACTI, or a custom flow.
  • Experience with FPGA prototyping or hardware emulation.

Responsibilities

  • Bring up inference workloads, including transformers (dense and MoE), attention and KV cache, expert routing, quantization, hybrid/SSM models, and retrieval, speech, vision, and recommendation workloads.
  • Bind Hugging Face and PyTorch workloads to the programming model and runtime, and execute them on the functional model for consistent software and architecture behavior.
  • Co-design tiling, scheduling, ISA, SRAM and HBM hierarchy, NoC traffic, and multi-chip mapping across various parallelism strategies.
  • Perform roofline and limiter analysis, and conduct design space exploration to identify and resolve bottlenecks between compiler and hardware views.
  • Develop Python energy and latency models using NumPy, Pandas, and Matplotlib for operators, tiling, memory traffic, and optical compute units.
  • Implement bit-accurate C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM for pre-tape-out software bring-up.
  • Contribute to the C++ event-driven simulation kernel, including coroutines, timed components, and traces.
  • Implement cycle-approximate and cycle-accurate performance, power, and area (PPA) models and align them with RTL through Verilator, SystemVerilog, and co-simulation.
  • Ensure consistency of numbers across different modeling and simulation approaches for the same workload and document discrepancies.
  • Define the modeling methodology for workload areas, including fidelity levels and arbitration for model disagreements.
  • Maintain the interface and register specifications as the source of truth for generating C++ and SystemVerilog views.
  • Mentor engineers developing models within your designated workload area.

Benefits

  • 100% coverage of base health plan premiums for you and your dependents
  • HSA contributions
  • Unlimited PTO
  • 401(k) matching
  • Stock option opportunities
  • Full suite of voluntary benefits, including Dental, Vision, Life, Hospital, Critical Illness, and Accident insurance.
  • Personalized Benefits: Choose plans that fit your life and take cash back for those that don’t.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service