product engineer, agent

Judgment LabsSan Francisco, CA

About The Position

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. The company ingests everything agents do in production: traces, tool calls, decisions, outcomes. Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals. Teams close the loop, shipping agent improvements validated against real production evidence. The Product Engineer will build the product experiences that make this loop legible, and will build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.

Requirements

  • Experience building and scaling end-to-end production systems, from data layer to UI
  • Strong technical problem-solving skills, especially in fast-changing, ambiguous environments
  • A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship
  • Hands-on experience building with LLMs or agents, or the drive to get there fast
  • Comfort working directly with customers to understand their needs and solve real-world problems
  • Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences

Responsibilities

  • Build the Judgment Agent to run large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.
  • Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes.
  • Design agent investigation interfaces to help engineers understand what their agents did and why, including long traces, tool calls, decisions, and failures. Debugging in the context of a reasoning loop and making thousand-step trajectories legible in minutes.
  • Design the Swarm UX to allow humans to watch a swarm work, redirect investigators, and consume findings without reading a hundred reports.
  • Build the workflows that turn production trajectories into datasets, judges, and regression checks, creating a seamless path from problem identification to fix verification.
  • Develop the platform infrastructure including workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.
  • Create an SDK and terminal-first experience for Judgment, allowing it to be summoned as a subagent mid-development in environments like Claude Code, Codex, and OpenCode sessions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service