The Interpretability team at Anthropic works to understand what's actually happening inside trained models and applies techniques to keep frontier AI safe as it rapidly improves. This role is an early hire on a new infrastructure effort within Interpretability, helping to define its charter. The job is to build the paved path that makes deep model access secure by default, private by design, and low-friction for every researcher. The work spans four areas: Security (design secure-by-default environments and access patterns), Privacy (build data-access patterns that ensure policy adherence), Data & Compute Management (manage research data at petabyte scale and make efficient use of large accelerator fleets), and Developer experience (agentic engineering, tooling and observability that keep researchers moving fast). In this role, you’ll be deeply embedded alongside Interpretability Researchers to understand their workflows, building your understanding of the research as you go. You’ll also bridge communication with Anthropic’s wider platform and security teams. Every hour of researcher friction you remove is multiplied across the whole organization, and the infrastructure you build sets the pace at which interpretability results reach real safety decisions.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior