Plato is an applied research lab building the foundational infrastructure to train specialized AI agents. We turn real-world data streams into high-fidelity simulated environments that generate the training signal needed to make capable models. Today, only a handful of players can train models for capable work. Compute and algorithms are rapidly commoditizing, but reinforcement learning data remains the bottleneck. Plato is changing that by automatically scaling training environments from proprietary real-world data. Our work supports frontier labs, hyperscalers, and enterprises building AI systems for complex, high-stakes work. Why This Role Matters Infrastructure is central to Plato's product and research loop. Generic cloud systems are not designed for long-running RL environments, persistent agent workspaces, replayable rollouts, storage-efficient forks, or recursive debugging loops. To train useful agents, we need infrastructure that makes environment construction, experimentation, evaluation, and iteration feel like one seamless system. As a Member of Technical Staff, Infrastructure / DevOps, you will own the systems that make Plato's research and training loops reliable at scale. Role Description You will build and operate the infrastructure behind long-horizon agent experiments, including environment VMs, storage-efficient snapshots and forks, orchestration for parallel agent fleets, shared workspaces, verifier workers, telemetry pipelines, deployment systems, and the operational tooling that lets researchers run thousands of experiments without thinking about the machinery underneath. This is not conventional cloud plumbing. You will be building infrastructure that directly shapes the quality, speed, and reliability of Plato's research.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed