Data Quality & Annotation Lead

Nimble RoboticsSan Francisco, CA
$100,000 - $120,000Onsite

About The Position

Nimble Brain turns real-world operational data into the training sets, evaluations, and feedback loops that make our robots smarter every day. The fuel for that engine is human demonstration data — skilled operators across our sites performing and recording the tasks our superhumanoids learn from — and the quality of every episode we collect sets the ceiling on every model we train. We're looking for a Data Quality & Annotation Lead to become the first dedicated owner of that quality bar. Own the rubric, not just run it: you'll write the acceptance criteria that define what a "good" episode is for every task family, build the audit workflows that enforce them as we scale from ~15 to 100+ operators across four sites this year, and stand up the annotation engine — including a remote annotation team turning around episode review overnight — that keeps labeled, trusted data flowing to research on schedule. This is a hands-on, metrics-driven, build-from-scratch role. You'll be based at our San Francisco HQ and spend heavy time at our collection sites, especially during operator ramps. Success looks like an episode acceptance rate that holds steady while the operator base grows 10x — with rubrics, audits, and dashboards that run like clockwork instead of heroics.

Requirements

  • 3+ years in data operations, annotation/labeling operations, or data collection QA for ML systems — robotics, autonomy, or teleoperation data strongly preferred.
  • A track record of building quality standards from scratch — rubrics, SOPs, QC workflows — not just executing against existing ones. Be ready to walk us through one you built.
  • Experience managing annotation or review teams, including remote/offshore or vendor/BPO teams, and driving their performance against SLAs.
  • Experience delivering direct, frequent quality feedback to operators or annotators — including the hard conversations when someone isn't meeting the bar.
  • Data fluency: able to build and own your own reporting and pressure-test the numbers — SQL, BI tools, or AI-assisted, we don't care how — rather than waiting on someone else.
  • Meticulous judgment on edge cases, paired with the pragmatism to ship a v1 rubric this week instead of a perfect one next quarter.
  • High agency and comfort with ambiguity in a fast-paced, high-growth environment — the org will triple around you this year.
  • Able to work in person out of our San Francisco HQ, with regular time at our collection sites.
  • Alignment with Nimble's values: relentlessly resourceful, humble, dependable, and committed to legendary impact.

Nice To Haves

  • Familiarity with imitation learning / robot learning data — what makes a demonstration usable for model training.
  • Experience scaling annotation or rater teams: pods, shift leads, inter-rater reliability programs.
  • Experience selecting and managing offshore annotation vendors, or standing up direct offshore teams (EOR, timezone-shifted workflows).
  • Multi-site operations experience.

Responsibilities

  • Own episode acceptance criteria for every task family — translate research and engineering data needs into operational, auditable standards, and version them as model training needs evolve.
  • Take over the annotation and episode-review queue in your first weeks, then build the team and workflows that scale it far beyond yourself.
  • Stand up and manage a remote annotation team (likely Philippines-based; direct hires or vendor/BPO) with an overnight turnaround SLA — episodes collected today are reviewed and scored before the next shift starts.
  • Design and run a sampling-based QA audit program with explicit coverage targets, including inter-rater reliability checks that keep annotators and auditors calibrated.
  • Run the operator quality feedback loop: operator-level quality scorecards delivered to site supervisors within 24 hours. You own the standard and the signal; supervisors own the coaching and people decisions.
  • Instrument your function: define the quality metrics (episode acceptance rate, audit coverage, feedback latency, operator quality distribution, annotation throughput) and build the operational dashboards your team runs on, partnering with our analytics function, which independently owns org-wide reporting.
  • Own the certification bar for new operator onboarding — no one collects production data without meeting it — while site teams run the day-to-day training reps.
  • Drive a standing weekly loop with research and engineering on failure modes, task-spec drift, and what "good" needs to mean next.

Benefits

  • Unlimited Flexible Time Off
  • Health Insurance
  • Paid Parental Leave
  • Commuter Benefits
  • Referral Bonus
  • 401k
  • Equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service