About The Position

G2 is the world's largest and most trusted software marketplace, now expanded with Capterra, SoftwareAdvice, and GetApp to become the largest source of online data and software insights. The company aims to transform the global B2B software industry by becoming the most trusted data foundation for buyers and sellers of software in the age of AI. G2 operates on PEAK values (Performance + Entrepreneurship + Authenticity + Kindness) and fosters a global, diverse community with people-led ERGs and a commitment to DEI and philanthropic work. This role focuses on the delivery of useful and credible evaluations of agents offered by software vendors, enhancing the system for validating agent performance and leading the core team behind it. The position will leverage "agent-first" ways of working and agentic engineering techniques for production delivery.

Requirements

  • 10+ years of professional programming experience in backend or full-stack environments
  • 2+ years of experience directly managing engineers
  • Expert-level proficiency in backend development using languages like Python, Java/Kotlin, Typescript/Javascript or Go; strong proficiency in associated backend frameworks like FastAPI or Node.js
  • Direct experience creating evaluations (evals) against a customer-facing agent, leveraging agent trajectory trace data and rubrics to measure an agent’s task completion rate, accuracy, correctness, or policy adherence
  • Direct experience using frontier models from OpenAI, Anthropic, or Google in LLM-as-a-judge applications
  • Regular use of coding agent harnesses like Claude Code, Codex, Opencode, or Pi as part of the daily development workflow
  • BS/BA degree in related field

Nice To Haves

  • Experience with agent tool use via direct integration or as mediated by MCP servers
  • Knowledge of STATE-Bench, tau2-bench or similar benchmark frameworks
  • Experience using Playwright, browser-use, Chrome DevTools MCP or others to drive user flows in the browser
  • Experience with durable execution frameworks & workflow solutions, like Temporal, DBOS, Cloudflare / Vercel Workflows, or similar; or experience with agent sandboxing using AWS E2B, Daytona, Cloudflare / Vercel Containers or similar

Responsibilities

  • Build and enhance agent evaluations that run against live agents.
  • Help scope out the feasibility and effort involved in evaluating an agent offered by a software vendor, identifying the most practical path to reaching a credible evaluation.
  • Own the full stack of the core evaluation system, from admin & API surface to eval workflow dispatch.
  • Design evaluation system primitives that generalize across software verticals, so onboarding a new category benefits from reuse.
  • Distill the appropriate and repeatable processes in onboarding and maintaining integrations into repeatable AI skills or agents to increase evaluation velocity.
  • Keep tabs on emerging frameworks and techniques for agent evaluation which offer improvements to evaluations, and provide industry-accepted means to distribute eval results.
  • Collaborate with data science peers in their pursuit to develop rich and proprietary benchmarks.
  • Facilitate the growth and mentorship of a core team of engineers, and be a source of technical leadership and subject matter expertise in evals for them.
  • Help socialize the use of evals in agent-oriented products built by the wider organization.

Benefits

  • G2 Gives program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service