Anthropic's Safeguards organization is responsible for developing policies, evaluations, and enforcement systems to prevent AI models from causing catastrophic harm. This role involves managing a research engineering team focused on biological safety. The team's work includes designing and executing capability evaluations for advanced AI models, curating training data for safety classifiers, training and refining these classifiers in collaboration with ML engineers, and assessing their performance against adversarial pressures in live traffic. The manager will define the technical strategy, prioritize team efforts, and be accountable for outcomes. This is a hands-on management position requiring significant time dedicated to team growth and direction, while maintaining enough technical expertise to review evaluation designs, analyze classifier failure modes, and effectively communicate with Research, Product, and Policy teams. The team's primary challenge is balancing the need for robust safeguards against sophisticated actors with the goal of not impeding legitimate research by the broader community using AI for life sciences. This tradeoff is an empirical problem that the team will measure and address.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Manager