About The Position

Plato is an applied research lab building the foundational infrastructure to train specialized AI agents. We turn real-world data streams into high-fidelity simulated environments that generate the training signal needed to make capable models. Our work supports frontier labs, hyperscalers, and enterprises building AI systems for complex, high-stakes work. Research engineering is central to Plato's product and research loop. The hard part of training specialized agents is not producing tasks that look plausible. It is finding tasks that are grounded in real workflows, difficult for current models, resistant to reward hacking, and useful as training signal. To do that, we need research engineers who can turn messy traces, model failures, and researcher hypotheses into environments, verifiers, rewards, evaluations, and curricula that improve continuously. As a Member of Technical Staff, Research Engineer, you will own the loop that discovers high-signal training targets for frontier models.

Requirements

  • Strong implementation ability and can turn ambiguous research ideas into working systems.
  • Experience with RL, LLM agents, computer-use agents, evals, post-training, synthetic data, simulation, or model behavior analysis.
  • Care deeply about whether a task is grounded, difficult, reward-hack-resistant, and capable of producing actual learning signal.
  • Comfortable interpreting ambiguous model behavior and negative results.
  • Enjoy building continuous research loops rather than static benchmark artifacts.

Responsibilities

  • Design experiments, build task generation systems, run evaluations, inspect model failures, and develop methods for mining tasks that are just out of reach of today's agents.
  • Consume real-world trajectories or researcher hypotheses, materialize realistic data, propose candidate tasks, benchmark those tasks against frontier computer-use and agent models, and hill-climb until you find the failures that produce useful learning signal.
  • Own the full loop: hypothesis, implementation, evaluation, analysis, iteration, and productionalization.
  • Discover model failure modes from real-world traces, agent telemetry, targeted researcher hypotheses, and customer workflows.
  • Generate realistic curricula grounded in actual workflows rather than toy synthetic benchmarks.
  • Benchmark candidate tasks against frontier CUA and agent models using pass rates, rollouts, and behavioral traces as difficulty signals.
  • Build hill-climbing loops that mutate, filter, and rescore tasks until they surface high-signal targets.
  • Study reward hackability, distribution mismatch, task realism, long-horizon failures, and transfer from simulation to deployed agents.
  • Turn research prototypes into reliable internal systems for continuous curriculum generation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service