Member of Technical Staff - Inference Systems

Liquid AICambridge, MA
Hybrid

About The Position

Spun out of MIT CSAIL, Liquid AI builds general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. They partner with enterprises across consumer electronics, automotive, life sciences, and financial services. The company is scaling rapidly and seeking exceptional people to join their team. This role is a core part of the team responsible for the engine layer that runs AI models in production and in partner environments, as well as for the benchmarking infrastructure used to evaluate their work and verify partner contributions. The position involves working closely with research and product teams, as well as directly with external engineering teams.

Requirements

  • Hands-on experience with at least one inference framework like llama.cpp, ONNX Runtime, or MLX, going beyond basic usage into internals and modification.
  • Experience designing and building benchmarking pipelines, including methodology, validation, and reproducibility.
  • Strong C++ and Python in performance-sensitive contexts.
  • Solid understanding of inference fundamentals: quantization, decoding strategies, memory layout, and how they interact.
  • Ability to pick up unfamiliar tools quickly and assess their utility.
  • Ability to design AI benchmarks and maintain high methodological standards.
  • Attention to inference details, understanding of tradeoffs, and thoroughness in verifying changes.
  • Commitment to proving output correctness for model ports.

Nice To Haves

  • Experience porting models across runtimes and verifying numerical correctness.
  • Prior work with external partners or clients in a technical validation or evaluation capacity.
  • Familiarity with edge inference targets and their constraints.

Responsibilities

  • Design and build benchmark suites that cover inference performance, model quality, and knowledge evaluation across different hardware targets.
  • Run external partner verifications: evaluate their solutions against our benchmarks, identify gaps, and clearly deliver findings.
  • Port models like LFM2 onto different runtimes and frameworks, and verify correctness end-to-end.
  • Maintain and extend the inference engine layer built on llama.cpp, ONNX, and MLX as new model architectures emerge from research.
  • Make benchmark results explainable and verifiable, so internal teams and partners can trust and reproduce them independently.

Benefits

  • Competitive base salary with equity in a unicorn-stage company
  • 100% of medical, dental, and vision premiums paid for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO
  • Company-wide Refill Days
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service