Runtime Engineer

MatXMountain View, CA
Hybrid

About The Position

MatX is building custom silicon for large-language-model inference and training, with HW/SW co-design across ISA, RTL, simulator, compiler, and kernels so each layer benefits from the others. The runtime owns the host-side stack and the contracts that bind those teams together.

Requirements

  • Strong experience in a systems programming language — Rust, C, C++, or Go — including memory management, allocator design, and FFI/ABI work
  • Have built Python interop layers in production (PyO3, ctypes, pybind11, or equivalent C-ABI bridging)
  • Have designed and maintained API or ABI contracts between teams — versioning, evolution, breaking-change discipline — not just consumed someone else's
  • Hands-on with at least one accelerator programming model (CUDA, ROCm, oneAPI Level Zero, TPU, or comparable) — enough to reason about device memory, async execution, and kernel launch
  • ML-systems literate — comfortable with the training and inference loop, what collectives do, what a tensor layout is. Research depth not required.

Nice To Haves

  • LLM inference internals — vLLM, TensorRT-LLM, or SGLang (paged attention, scheduler design)
  • Rust at depth, including proc macros, unsafe with soundness reasoning, and complex lifetime/trait work
  • Custom allocator design (slab, paged, arena) or other low-level memory work
  • ML framework integration experience (PyTorch custom backends, JAX/XLA, ONNX runtime)
  • Profiler or tracing infrastructure work (perfetto, Nsight, or a custom stack)
  • Driver-adjacent or kernel-bypass work, or prior new-silicon bring-up

Responsibilities

  • Build the host-side interface library — device memory management, DMA, streams and events, sync primitives — that every compiler-emitted program runs on top of
  • Own and extend the executable format: the compiler→runtime contract, its versioning, the weight and quantization layouts that let compiler and runtime evolve independently
  • Design the custom-kernel ABI — calling convention, sync semantics, lifecycle — and the host-side marshaling layer (DLPack, the buffer protocol, numpy) that gets Python tensors to the device
  • Build Python bindings via PyO3, with a C-ABI shim as the alternative integration path for downstream consumers
  • Build the LLM inference serving stack — paged KV cache, continuous batching, request scheduling, token streaming — and the cluster orchestration primitives underneath it
  • Bring up interconnect topology from the host and own the failure-detection and clean-teardown path for stop-restructure-resume recovery across racks
  • Design what the chip exposes to host-side profilers and debuggers — perf counters, traces, and the Python surfaces ML engineers actually use — and hit measurable performance targets on runtime overhead and serving throughput

Benefits

  • 4 weeks PTO (accrued)
  • 12 company Holidays
  • up to 3 weeks remote work
  • Company-subsidized Medical (Kaiser or Anthem) for employees & dependents
  • Guardian Dental and Vision insurances for employee & dependents
  • life insurance (employee only)
  • HSA and FSA offerings via Lively
  • Roth IRA/ 401K (or both) retirement plans
  • up to 5% company contribution to 401K
  • 100% company-paid life insurance (up to $300K)
  • long-term disability insurances
  • $1500 Professional Development Budget (per year)
  • Onsite team lunch & dinner Monday - Friday
  • Company Uber account for commute
  • Reimbursement for train rides
  • $50/mo to use on the perk you value most
  • $35/mo for cellular
  • $40/mo for wifi
  • 100% paid mental health benefit via SpringHealth and Guardian EAP
  • Up to 12 weeks paid parental leave regardless of path to parenthood
  • 10 weeks pregnancy disability leave
  • flexible return-to-work hours
  • Benepass reproductive health & parental benefit
  • Up to $20K/month for AI Resources
  • Dedicated internal AI Tooling Team
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service