Member of Technical Staff - Research

CruciblSan Francisco, CA
Hybrid

About The Position

Crucibl Management consulting is a $400B industry built on selling intelligence-for-hire. The model is breaking - not slowly, but now. The biggest frontier in enterprise AI isn't better models. It's judgment at scale: taking the messy, ambiguous decisions that run 85% of the global economy and making them faster, crisper, and more defensible. Crucibl is building that. We're profitable, growing fast, and delivering for Fortune 500 clients. We raised a $10M seed from Tier 1 VCs - then kept growing on revenue. The client pipeline is full. We're scaling to meet it.

Requirements

  • At least one undeniable signal of excellence - published research at a top venue, research role at a frontier lab, or a track record of novel technical contributions
  • Deep understanding of how LLMs reason, fail, and can be evaluated - this isn't a theoretical interest, it's core to the job
  • Strong fundamentals in ML, statistics, or NLP - we're too small for hand-holding on the technical side
  • Comfortable moving from open-ended research question to a testable hypothesis quickly

Nice To Haves

  • Experience designing evaluation frameworks or benchmarks for reasoning, decision-making, or agentic systems
  • Published or shipped work on uncertainty, calibration, multi-step reasoning, or LLM evaluation
  • Has operated in a high-growth, high-ambiguity environment
  • Thinks like an owner - big picture, not just your part

Responsibilities

  • Research how frontier models reason through ambiguous, high-stakes business decisions - and where they fail
  • Design novel methods for reasoning, evaluation, and calibration that go beyond standard benchmarks
  • Translate open problems in reasoning, uncertainty, and multi-step decision-making into approaches we can test and ship
  • Define what "good judgment" looks like for a model, and build the evaluation frameworks to measure it
  • Design experiments that reveal real failure modes, not just leaderboard scores
  • Turn findings into concrete recommendations for the product and applied teams
  • Shape technical vision and roadmap alongside the founding team
  • Bring outside research thinking into a company solving a problem few labs are focused on
  • Define what rigorous, applied research looks like in an AI-first organization
  • Raise the bar for the team as it grows

Benefits

  • Meaningful Equity: Every offer includes a comprehensive salary and equity package.
  • Hybrid Schedule: 3 days in-office with real flexibility around the rest. We care about output, not optics.
  • Full Health Coverage: Medical, dental, and vision.
  • Daily Lunch & Snacks: Fueled and focused, on us.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service