Research Scientist

AugerBellevue, WA

About The Position

The core of this work is training foundational models for supply chain expertise, not wrapping a generalist frontier model in a better prompt. We think a specialist model, trained deep on the domain, beats a generalist model on the problems that actually matter here, and gets us to a level of inference speed, cost, and reliability that routing every decision through a frontier API simply can't reach. What Makes You Succeed Here We'd rather see your checkpoint get quantized, distilled, and forked into someone else's production stack than see it top a leaderboard for a week and disappear. A few things from The Auger Edge show up again and again in the people who do well here. The instinct to Explore to Evolve looks like this in practice: you don't just call .fit() on a technique, you can derive why it works, and you'll rebuild the pipeline from the tokenizer up when the domain demands it, whether that's continued pretraining into a knowledge-intensive vertical or an eval harness that measures something real instead of something convenient. If you've built evaluation frameworks specifically to catch what standard benchmarks miss, you're already living Own the Fall, Rise Stronger: you treat a bad eval run as signal, not shame, and the loop from "here's where it breaks" to "here's the next checkpoint" is short. The field dresses complexity up as sophistication constantly, which is exactly what it means to Crush Complexity here: we want the person who ships the clean dataset and the clean eval that a teammate can pick up cold, not the clever bespoke pipeline only its author can operate. Tech-leading through v1, v2, v3, each release measurably stronger than the last, is what All In, All the Time looks like day to day, and it's also why your job isn't done at a passing eval or a merged PR. It's done when you've watched the checkpoint run flawlessly in production, under real load, on real customer data. Ask anyone who's been here a while what that means in practice: the job is never actually done, there's always a v4. What You Bring You've built foundational training data at scale, corpora and not just models, and understand that what goes into a model matters as much as its architecture. You've led a project across multiple release cycles, each one measurably better than the last. You've designed evaluation methodology that goes beyond standard benchmarks, built specifically to surface what those benchmarks miss. You've adapted general purpose models to specialized, knowledge intensive domains and understand what actually transfers versus what has to be rebuilt. You've created datasets that other researchers and practitioners now build on. You've taken research past the paper and into a real, end to end system that people other than researchers actually use. Recognition, best paper or outstanding paper or otherwise, has followed the work, but wasn't the point of the work. We're not hiring for a specific problem or a specific product. We're hiring for a pattern. If you read that list and thought "yes, and also," we want to talk to you.

Requirements

  • Built foundational training data at scale, corpora and not just models.
  • Understands that what goes into a model matters as much as its architecture.
  • Led a project across multiple release cycles, each one measurably better than the last.
  • Designed evaluation methodology that goes beyond standard benchmarks, built specifically to surface what those benchmarks miss.
  • Adapted general purpose models to specialized, knowledge intensive domains and understands what actually transfers versus what has to be rebuilt.
  • Created datasets that other researchers and practitioners now build on.
  • Taken research past the paper and into a real, end to end system that people other than researchers actually use.
  • Work has been recognized (best paper or outstanding paper or otherwise), but recognition wasn't the point of the work.

Nice To Haves

  • Checkpoint quantized, distilled, and forked into someone else's production stack.
  • Instinct to Explore to Evolve.
  • Experience with continued pretraining into a knowledge-intensive vertical.
  • Experience with an eval harness that measures something real instead of something convenient.
  • Experience with Own the Fall, Rise Stronger mindset (treating bad eval runs as signal, short loop from 'here's where it breaks' to 'here's the next checkpoint').
  • Experience with Crush Complexity (shipping clean datasets and evals that teammates can pick up cold).
  • Experience with All In, All the Time (tech-leading through v1, v2, v3, and ensuring job is done when checkpoint runs flawlessly in production).

Responsibilities

  • Training foundational models for supply chain expertise.
  • Deriving why techniques work and rebuilding pipelines from the tokenizer up when the domain demands it.
  • Building evaluation frameworks to catch what standard benchmarks miss.
  • Adapting general purpose models to specialized, knowledge intensive domains.
  • Taking research past the paper and into a real, end to end system.
  • Ensuring checkpoints run flawlessly in production, under real load, on real customer data.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service