Member of Technical Staff, Head of Quality

PlatoSan Francisco, CA
Onsite

About The Position

Plato is an applied research lab building the environments used to train specialized AI agents. We turn proprietary real-world data into high-fidelity simulations that produce the dense reinforcement learning signal required to train frontier models. Compute and baseline architectures are rapidly commoditizing; environment design and RL data are the true bottlenecks. Today, models don’t fail from a lack of compute—they fail because bad task design, leaky reward functions, and brittle verifiers train them to cheat instead of learn. Plato exists to make RL signal provable, grounded, and robust at scale. We’re based in San Francisco and backed by leading investors and researchers across top frontier labs. Why Apply Own the Ground Truth: You will own the final checkpoint between raw environments and customer model runs. If an environment teaches a model the wrong behavior, you pull the plug. Massive Leverage: We don’t solve quality by throwing armies of manual labelers at a spreadsheet. You will architect the automated judge harnesses, red-teaming agents, and telemetry that enforce quality programmatically. Founding Impact: As Head of Quality (Member of Technical Staff), you will build Plato’s QA systems from scratch, set the technical bar, and scale the team. Frontier Signal: Work directly on the failure modes frontier labs face when scaling post-training, test-time compute, and agentic workflows. The Role In RL, quality is not a polish step—it is the training signal itself. When task designs are ambiguous, sandboxes lack proper isolation, or reward functions contain loopholes, agents do what optimizers always do: exploit the grader. A broken environment doesn’t just burn cluster hours; it actively poisons downstream model weights. As Head of Quality, you will own the standard for what constitutes real learning signal across every task, verifier, and environment Plato ships. You’ll be the adversarial mind finding the loopholes before the model does, the engineer automating the test suites, and the leader directing the team running verification. What You’ll Do Gate Final Delivery: Own the final sign-off before environments and datasets ship to frontier labs, auditing trajectories, tasks, and reward dynamics for correctness, feasibility, and signal density. Harden Verifiers & Tasks: Build automated red-teaming suites to stress-test task feasibility and verifier integrity, aggressively eliminating reward hacking, grader tampering, and impossible task traps. Automate Verification Infra: Architect judge models, sandbox replay harnesses, and rollout forensics to catch synthetic drift, out-of-distribution behaviors, and leaky states upstream. Build & Lead the Verification Team: Hire and direct a high-agency team of QA engineers, domain specialists, and technical reviewers, blending automated agentic checks with deep human-in-the-loop review. Close the Loop with Research: Translate downstream model failure modes into concrete generator constraints so defects are prevented at the generation stage rather than caught in review. Introduction Plato is an applied research lab building the foundational infrastructure to train specialized AI agents. We turn real-world data streams into high-fidelity simulated environments that generate the training signal needed to make capable models. Our work supports frontier labs, hyperscalers, and enterprises building AI systems for complex, high-stakes work. Today, only a handful of players can train models for capable work. Compute and algorithms are rapidly commoditizing, but reinforcement learning data remains the bottleneck. Plato is changing that by automatically scaling training environments from proprietary real-world data.

Requirements

  • 3+ years of experience in software engineering, ML engineering, or research systems, with proficiency in modern languages (Python, etc.).
  • You understand how optimizers exploit edges. You naturally think about how an agent could game a reward function, break a sandbox, or fake completion.
  • Proven ability to dive into raw rollouts, agent reasoning traces, and verification code to spot subtle ungrounded assumptions or hallucinated logic.
  • Experience building or leading a technical evaluation, QA, or data verification function from zero to one.
  • Comfort holding the line on delivery under intense customer pressure. You understand that shipping contaminated data is far worse than shipping late.

Nice To Haves

  • Experience with RL training dynamics, automated LLM evals, agent sandboxing, or synthetic trajectory generation.
  • Background designing adversarial test suites, code execution verifiers, or formal verification systems.
  • Experience handling client-facing technical evaluations and failure postmortems with frontier AI research teams.

Responsibilities

  • Own the final sign-off before environments and datasets ship to frontier labs, auditing trajectories, tasks, and reward dynamics for correctness, feasibility, and signal density.
  • Build automated red-teaming suites to stress-test task feasibility and verifier integrity, aggressively eliminating reward hacking, grader tampering, and impossible task traps.
  • Architect judge models, sandbox replay harnesses, and rollout forensics to catch synthetic drift, out-of-distribution behaviors, and leaky states upstream.
  • Hire and direct a high-agency team of QA engineers, domain specialists, and technical reviewers, blending automated agentic checks with deep human-in-the-loop review.
  • Translate downstream model failure modes into concrete generator constraints so defects are prevented at the generation stage rather than caught in review.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service