Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence. We are the team behind some of the earliest and most influential research in AI evaluation like FinanceBench, Lynx, SimpleSafetyTests, CopyrightCatcher, Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more. The AI Platform Engineer sits between research and engineering, turning research workflows into services that other teams can self-serve. The work centres on building and operating the several internal platforms that the research and engineering teams rely on daily. These are production products with real users — backends, dashboards, CLIs and SDKs — built for colleagues inside the company rather than for an abstract audience. The second half of the role is the foundation those platforms stand on: deploying models and keeping them served, both on managed inference providers and on internally operated GPUs, together with the cluster services that make GPU capacity self-serve rather than a matter of negotiation. This is a hands-on, delivery-oriented platform position rather than a research one: the emphasis falls on the systems and tooling that make other teams' work possible. Model training is part of the surrounding environment and remains accessible, but it is not the focus of the position. In this role, you will: Building and operating the internal platforms end to end — backends, storage, dashboards, and the CLI and SDK surfaces they are driven through — including multi-tenancy, sign-in and access control. Owning workload orchestration on the GPU clusters and the services around them — submission and scheduling, provisioning, quotas and placement. Deploying models and keeping them served, on managed inference providers and on internally operated GPUs — sizing each deployment for its hardware, and writing serving wrappers where no off-the-shelf engine fits. Building the evaluation surface, so that a result stays comparable across months, colleagues and models. Establishing and hardening CI/CD, containerized workflows and release safety. Instrumenting the platform with logging, metrics and alerting, so that divergence between what a service promises and what it serves is caught by a test rather than by a customer. Adding agent surfaces to the platform, with whatever an agent resolves written back as auditable configuration. Partnering with research engineers to productionize their experiments, and carrying production issues through to resolution — including on-call for systems built in this role.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior