Staff AI Evaluation Lead

FetchRemote,
$129,000 - $152,000Hybrid

About The Position

At Fetch, we’re building AI and automation systems that make our work smarter, faster, and more scalable. The AI Operations team ensures our models, automations, and LLM systems perform with quality, reliability, and measurable impact. As a Staff AI Evaluation Lead, you’ll own automation and evaluation programs across AI Operations. You’ll translate business and functional goals into scalable systems, define how we measure quality, and ensure automation and evaluation become durable, high-impact capabilities across the organization. This role is ideal for someone who combines deep technical problem-solving with systems-level thinking and strong cross-functional leadership. This is a full-time role that can be held from one of our US offices or remotely in the United States.

Requirements

  • 8+ years of professional experience AI, machine learning, operational automations, or a related field.
  • Proven ability to lead complex automation or evaluation initiatives across systems or teams
  • Experience designing, building, and scaling evaluation frameworks, datasets, and quality systems at scale for LLMs or AI products
  • Fluency in SQL, JSON, APIs, and scripting with AI assistance, and ownership of the technical direction of the evaluation stack
  • Deep working knowledge of LLM and agentic system behavior, automation platforms, and system design, demonstrated in systems you have built
  • Experience with data pipelines, APIs, and production systems
  • Experience in influencing cross-functional stakeholders and aligning priorities
  • Demonstrated ability to define metrics and drive measurable business impact

Nice To Haves

  • Experience mentoring or leading technical contributors
  • Experience evaluating agentic systems in production at scale
  • Experience setting technical standards adopted across an organization

Responsibilities

  • Own programs: Lead complex, high-impact automation and evaluation initiatives across workflows or teams
  • Design scalable solutions: Architect end-to-end workflows integrating datasets, evaluations, automations, and HITL processes; build the harness and reusable components other analysts build against; build evaluation pipelines that run against production without hands-on operation
  • Establish standards: Define dataset standards, evaluation methodology, failure taxonomy, and quality measurement across AI Operations; verify that projects you don't run are meeting them; own the process by which those standards get made and used
  • Own metric definitions: Create and maintain the metric definitions the org evaluates against, validate them against real production behavior, and revise them as models, tooling, and system architectures change
  • Set the bar for production: Define the quality bar a system clears before it reaches production and where human review stays in the loop, and pull a system back when it stops meeting the bar
  • Allocate evaluation depth by risk: Decide which systems get what depth of evaluation given finite capacity, name the risk accepted on the rest, and make that tradeoff visible to project stakeholders
  • Keep the evaluation stack current: Define when a change to a model or platform requires re-baselining across the org and own the process for doing it; evaluate new models and AI capabilities as they ship and decide what the org adopts
  • Improve systems at scale: Lead redesign of workflows, tooling, and processes to improve performance and durability across AI Operations, not only within your own programs
  • Drive cross-functional alignment: Influence priorities and partner with Engineering, Product, and AI teams to deliver solutions
  • Measure and communicate impact: Define success metrics and communicate performance and recommendations to leadership
  • Elevate the team: Raise the bar through your work and actively mentor others by sharing approaches, guiding problem-solving, and enabling the team to build stronger automation and evaluation capabilities
  • Drive innovation: Stay current on emerging tools and approaches; pilot and translate them into actionable improvements for the team

Benefits

  • competitive compensation packages including base, equity, and benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service