Senior | Staff Software Engineer - AI / ML

Snorkel AI•San Francisco, CA
•$208,000 - $315,000

About The Position

Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI. You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it.

Requirements

  • 5+ years building production ML or software systems, with end-to-end ownership from prototype to production
  • Hands-on experience running LLM or ML workloads in production, and comfort reasoning about non-deterministic systems
  • Deep grounding in statistics and experimentation: experiment design, hypothesis testing, sampling, and confidence intervals
  • Strong Python and software engineering fundamentals, including testing, code review, and system design
  • Experience designing evaluations and interpreting results rigorously
  • A habit of finding high-impact problems before they are assigned, and clear communication with researchers, engineers, and business partners

Nice To Haves

  • Fine-tuning and serving open-weight models, and judging when a smaller model meets the quality bar
  • Building LLM evaluation or experimentation platforms, model gateways, or routing systems
  • Experience with agentic workloads, benchmarks, or RL environments
  • A record of taking research into production: publications, open-source work, or shipped research-driven features
  • MS or PhD in Computer Science, Machine Learning, Statistics, or a related field

Responsibilities

  • Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating.
  • AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution.
  • Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other parameter-efficient methods) where they match frontier quality, and know when they don't.
  • Predictive difficulty. Build models that estimate how hard a task is for frontier systems before running a single rollout.
  • Measurement for AI data. Build golden datasets, quantify the accuracy and calibration of LLM-as-judge systems, and make quality reproducible across projects.
  • Research to production. Turn research prototypes into reusable, configurable components that forward deployed engineers and researchers use on every project.

Benefits

  • Meaningful opportunities to shape priorities and initiatives
  • Influence key strategic decisions
  • Directly impact ongoing success
  • Deepen technical expertise
  • Explore leadership opportunities
  • Learn new skills across multiple functions
  • Support for building career in an environment designed for growth, learning, and shared success
  • Equal employment opportunities
  • Reasonable accommodation for individuals with disabilities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service