As Sweden's national center for applied AI, AI Sweden is looking for a master thesis student to join their team. This thesis will investigate whether machine unlearning genuinely erases target information or merely suppresses its output expression. Current evaluations often measure success behaviorally, treating information as forgotten if a model fails to produce the target answer under direct questioning. However, surface refusal is deceptive: deleted facts routinely persist in intermediate hidden states or resurface under adversarial prompting, steering, and fine-tuning. Without empirical stress-testing, unlearned models deployed in public risk regulatory exposure and data leakage. AI Sweden is leading the development of LeakPro, an open-source privacy auditing tool designed to assess information leakage risks in machine learning models. This initiative aims to evaluate the risk of sensitive information disclosure when models trained on confidential data are made publicly available. Existing benchmarks approximate the counterfactual with a single retained checkpoint and score behavior at rest, which cannot separate genuine unlearning effects from ordinary seed-to-seed training variance. This thesis asks if a multi-seed counterfactual audit can detect residual target knowledge that behavioral evaluation reports as successfully forgotten, and for which classes of unlearning method does this gap between genuine erasure and behavioral suppression appear.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Entry Level