Product Lead, Rippling AI – Agent Harness & Runtime

RipplingSan Francisco, CA
$174,000 - $290,000Onsite

About The Position

Model quality is converging across providers. The differentiator that remains — and compounds — is the layer around the model: how it holds context, recovers from tool failure, knows what it's allowed to touch, and gets measurably better over time without a human rewriting its instructions every week. That layer is the harness. It is the difference between a demo and a system you can put in front of a customer's payroll data. We're hiring the PM who owns that layer for Rippling's agents: our own harness, the memory systems that feed it, and the self-improvement loop that tunes it. This is not a PM role bolted onto an engineering team's backlog. You will make build-vs-buy calls on runtime architecture, define what "done" means for a memory system, and be the person who can sit in a design review with staff engineers and catch a bad abstraction before it ships.

Requirements

  • A former software engineer-turned-PM with 5+ years of combined experience.
  • You've shipped production systems yourself, not just specced them.
  • You've built or shipped agentic systems in production, not just prototyped them in a notebook.
  • You can speak precisely about the difference between an agent framework (the blueprint) and a harness (the runtime that actually executes and recovers), and you don't use the terms interchangeably.
  • You think in state machines and failure modes by default: what happens when the tool call times out, when the model hallucinates a tool that doesn't exist, when two agents claim the same lock.
  • You've debugged distributed systems before you ever wrote a spec for one.
  • You have a strong point of view on evals and can design a benchmark that predicts production behavior instead of rewarding overfit.
  • You're fluent in the concepts this space actually uses day to day: context window management, tool orchestration, RAG vs. semantic vs. episodic memory, OAuth grant types and blast radius, sandbox isolation, RL/fine-tuning vs. non-parametric adaptation.
  • You default to writing the design doc yourself when the team is moving too fast to wait for consensus, and you're just as comfortable being told your design is wrong by an engineer with more context.

Responsibilities

  • Inform the architecture that turns a model into a worker, including how it plans, calls tools, interprets results, and decides to continue, retry, or escalate. You'll define the contract between orchestration and the harness, any required guardrails, and how to refine the harness to support the agents we’re building across Rippling.
  • Working memory inside a single run, episodic memory across sessions, understanding and defining the boundary between what belongs in context versus what belongs in a tool call. You'll help define what gets persisted, what gets compacted, and what gets forgotten on purpose.
  • The feedback loop that makes the harness better without a human editing prompts by hand — eval-driven tuning, trajectory review, counterfactual benchmarking against prior versions. You own the definition of "improved": task completion rate, escalation-to-human rate, cost and latency per successfully completed task, and regression rate on the eval suite you help build.
  • Sandboxing and execution environments, permission tiers and connector-level ACLs, credential handling so the model never sees a raw secret, and the observability stack that lets an engineer answer "why did this agent do that" six weeks after the fact.
  • As the core harness matures, you'll extend it into computer-use and autonomous web/research capabilities — agents that navigate interfaces without an API, and multi-step research tasks that synthesize across sources.

Benefits

  • competitive salary + benefits + equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service