Anthropic's Safeguards team builds systems to detect and mitigate misuse of AI models, including individual policy violations and sophisticated, coordinated attacks. A significant part of this work involves lightweight detection methods trained on model internals, enabling cost-effective and scalable identification of harmful behavior. This role focuses on owning the infrastructure that supports this research, including tooling for experiments, training detection methods, and selecting detections for launch. It bridges the gap between research and production, ensuring fast iteration for researchers and reliable results for detection systems as models evolve. The engineer will tackle novel systems problems at scale, building abstractions, pipelines, and tooling to maintain research velocity amidst shifting requirements. The ideal candidate will have a proven ability to solve large-scale systems and data problems and a strong interest in developing deep machine learning expertise.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level