Modeling Architect

NeurophosAustin, TX
$170,000 - $200,000Onsite

About The Position

We are seeking a modeling architect for hands-on architecture modeling of the T100 optical inference accelerator, with hardware/software co-design in the loop. You will work alongside senior engineers across two tracks. The first is analytical and system performance: roofline and limiter analyses, architecture performance models, workload setup, and the resulting plots and reports. The second is hardware models: functional, performance, and power models of compute blocks, memory, and hardware/software interfaces built in an event-driven simulator, plus RTL simulation. We expect real depth in one track and will help you build breadth across both. Most engineers start with a bounded piece, a single workload, a hardware block, or one layer of the model stack, and take on the surrounding area as the models mature.

Requirements

  • BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
  • 3+ years of experience in hardware modeling, performance simulation, computer architecture, or related work.
  • Proficiency in Python or modern C++ (C++17 or later). Python-first and C++-first backgrounds are both welcome.
  • Working knowledge of computer architecture and microarchitecture, including pipelines, caches, memory hierarchies, and instruction set architecture (ISA).
  • Ability to turn an LLM, GEMM, or accelerator paper into a workload config using Hugging Face or PyTorch.
  • Strong debugging skills and the habit of writing down what was run, what was assumed, and what the number means.

Nice To Haves

  • MS or PhD in Computer Engineering, Electrical Engineering, or Computer Science.
  • Experience with roofline analysis, limiter analysis, GPU benchmarking, or correlating a model against published numbers.
  • Event-driven, cycle-approximate, or cycle-accurate simulation with SystemC, gem5, SST, or a custom kernel.
  • SystemVerilog, Verilog, Verilator, or RTL co-simulation.
  • Memory and interconnect experience with HBM, DRAM, cache, SRAM, network-on-chip (NoC), AXI, or DMA.
  • Familiarity with CUDA, GPU programming, or PyTorch internals.

Responsibilities

  • Bring up inference workloads as they ship, including dense and Mixture of Experts (MoE) transformers, hybrid/SSM models, prefill versus decode, KV cache, expert routing, and quantization, plus retrieval, speech, vision, and recommendation workloads where they map onto the accelerator.
  • Bind Hugging Face and PyTorch workloads to the programming model and run them on the functional model.
  • Co-design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, and multi-chip mapping.
  • Build in one or more layers of the modeling stack: roofline and limiter studies; Python energy and latency models; C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM; cycle-approximate performance and power models; and RTL simulation with Verilator and SystemVerilog.
  • Own the tests, configs, and plots behind a result so anyone can rerun it and see what was assumed.
  • Use coding agents on real multi-file edits, and own the review of the C++ and SystemVerilog they generate.
  • Share results with the architects setting the design, as well as the compiler, runtime, and RTL teams, and carry their questions back into the model stack.

Benefits

  • 100% coverage of base health plan premiums for you and your dependents, plus HSA contributions.
  • Unlimited PTO.
  • 401(k) matching and stock option opportunities.
  • Full suite of voluntary benefits, including Dental, Vision, Life, Hospital, Critical Illness, and Accident insurance.
  • Personalized Benefits. Choose the plans that fit your life and take the cash back for those that don’t.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service