Member of Technical Staff (ML Eval)

NemaSan Francisco, CA

About The Position

We're building AI agents that understand how safety-critical hardware gets built and we need your expertise. In this role you'll own the LLM evaluation infrastructure. You'll build evals that catch regressions before they impact customers, and you'll work directly with the product team to define what "good" looks like for Nema Agent-generated requirements, test cases, and documentation.

Requirements

  • A bias to action
  • Experience building evaluation infrastructure
  • The patience to label data yourself when needed

Nice To Haves

  • Experience evaluating LLMs in regulated industries (aerospace, defense, automotive, medical)

Responsibilities

  • Build and maintain evaluation pipelines for LLM outputs
  • Design domain-specific benchmarks for hardware engineering workflow
  • Instrument our AI features to collect human feedback and measure real-world performance
  • Run experiments comparing model architectures and context management strategies
  • Work with customers to understand failure modes and build evals that catch them
  • Do whatever it takes to help the company win
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service