This position develops and evaluates machine unlearning and model editing methods that selectively reduce hazardous biological capabilities in AI systems while preserving beneficial scientific functions. The researcher reports to Assistant Professor Tom Hartvigsen and will work closely with other SDS faculty members including Chirag Agarwal, and Stephen Turner, and works with faculty in interpretability and with a laboratory partner that leads adversarial red teaming. The role centers on implementing, innovating, and comparing model editing and unlearning methods, measuring safety--utility tradeoffs against both benchmarks and realistic task batteries, and leading technical development of an open evaluation suite for AI biosecurity. Strong familiarity with biology and biosecurity is important, as the work targets biological capabilities and connects to a human-subjects evaluation running in parallel.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
Ph.D. or professional degree