Anthropic is seeking Cyber Evaluations Engineers to develop and execute evaluations that assess the cyber-relevant capabilities and robustness of their AI models. This role involves designing new evaluation methods, conducting per-release robustness testing, and analyzing data related to jailbreaks and prompt bypasses to understand the effectiveness of safeguards. The engineer will also design probes to detect cyber abuse in production and collaborate with the policy team to shape the overall detection architecture. Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The company is a rapidly growing team of researchers, engineers, policy experts, and business leaders dedicated to this mission.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level