Modeling Architect

NeurophosAustin, TX
Onsite

About The Position

Neurophos is developing a novel AI inference accelerator using silicon photonics and an active, programmable metasurface to perform matrix multiplications at the speed of light. This approach aims to overcome the power limitations of traditional silicon by offering significantly higher energy efficiency and performance for large-scale AI inference. The company has secured $110M in Series A funding and is seeking a Modeling Architect to contribute to the architecture modeling of their T100 optical inference accelerator, with a focus on hardware/software co-design. The role involves both analytical system performance modeling and building hardware models within an event-driven simulator, alongside RTL simulation. The position is full-time and onsite, located in Austin, TX or Sunnyvale, CA.

Requirements

  • BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
  • 3+ years of experience in hardware modeling, performance simulation, computer architecture, or related work.
  • Proficiency in Python or modern C++ (C++17 or later).
  • Working knowledge of computer architecture and microarchitecture, including pipelines, caches, memory hierarchies, and instruction set architecture (ISA).
  • Ability to translate LLM, GEMM, or accelerator papers into workload configurations using Hugging Face or PyTorch.
  • Strong debugging skills and a methodical approach to documenting assumptions and results.

Nice To Haves

  • MS or PhD in Computer Engineering, Electrical Engineering, or Computer Science.
  • Experience with roofline analysis, limiter analysis, GPU benchmarking, or correlating models against published numbers.
  • Experience with event-driven, cycle-approximate, or cycle-accurate simulation using SystemC, gem5, SST, or a custom kernel.
  • Experience with SystemVerilog, Verilog, Verilator, or RTL co-simulation.
  • Memory and interconnect experience with HBM, DRAM, cache, SRAM, network-on-chip (NoC), AXI, or DMA.
  • Familiarity with CUDA, GPU programming, or PyTorch internals.

Responsibilities

  • Bring up inference workloads including transformers (dense and MoE), hybrid/SSM models, prefill vs. decode, KV cache, expert routing, quantization, retrieval, speech, vision, and recommendation workloads.
  • Bind Hugging Face and PyTorch workloads to the programming model and execute them on the functional model.
  • Co-design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, and multi-chip mapping.
  • Build one or more layers of the modeling stack: roofline and limiter studies; Python energy and latency models; C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM; cycle-approximate performance and power models; and RTL simulation with Verilator and SystemVerilog.
  • Own the tests, configurations, and plots for results to ensure reproducibility.
  • Utilize coding agents for multi-file edits and review the generated C++ and SystemVerilog code.
  • Share results with design architects, compiler, runtime, and RTL teams, and incorporate their feedback into the model stack.

Benefits

  • 100% coverage of base health plan premiums for you and your dependents.
  • HSA contributions.
  • Unlimited PTO.
  • 401(k) matching.
  • Stock option opportunities.
  • Full suite of voluntary benefits, including Dental, Vision, Life, Hospital, Critical Illness, and Accident insurance.
  • Personalized Benefits: Choose plans that fit your life and receive cash back for unused options.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service