About The Position

As Sweden's national center for applied AI, AI Sweden is looking for a master thesis student to join their team. This thesis will investigate whether machine unlearning genuinely erases target information or merely suppresses its output expression. Current evaluations often measure success behaviorally, treating information as forgotten if a model fails to produce the target answer under direct questioning. However, surface refusal is deceptive: deleted facts routinely persist in intermediate hidden states or resurface under adversarial prompting, steering, and fine-tuning. Without empirical stress-testing, unlearned models deployed in public risk regulatory exposure and data leakage. AI Sweden is leading the development of LeakPro, an open-source privacy auditing tool designed to assess information leakage risks in machine learning models. This initiative aims to evaluate the risk of sensitive information disclosure when models trained on confidential data are made publicly available. Existing benchmarks approximate the counterfactual with a single retained checkpoint and score behavior at rest, which cannot separate genuine unlearning effects from ordinary seed-to-seed training variance. This thesis asks if a multi-seed counterfactual audit can detect residual target knowledge that behavioral evaluation reports as successfully forgotten, and for which classes of unlearning method does this gap between genuine erasure and behavioral suppression appear.

Requirements

  • Ongoing Master’s studies in Computer Science, Data Science, Engineering Physics, Complex Adaptive Systems, Machine Learning, or a related field.
  • Comfortable with Python and deep learning.
  • Comfortable with the reality that an experiment might yield unexpected results.

Nice To Haves

  • Curious, independent, and self-driven MSc student.

Responsibilities

  • Literature study of LLM unlearning and auditing: Summarize (i) approximate unlearning methods, (ii) evaluating unlearning though model outputs, internal layer memory, and relearning speed (iii) sequential unlearning dynamics, and (iv) suitable datasets and benchmark architectures.
  • Implementation of a counterfactual audit benchmark: Train a multi-seed counterfactual reference distribution on target-omitted data. Evaluate unlearned models across behavioral leakage, layer-level representation probes, and knowledge recovery audits.
  • Longitudinal and sequential unlearning analysis: Audit model behavior across repeated independent, semantically clustered, and overlapping deletion requests to construct a threat-model-specific erasure-suppression profile.
  • Contribute to the open-source platform LeakPro by integrating the unlearning audit outputs and taxonomy into LeakPro's privacy framework (if time permits).

Benefits

  • Work alongside leading AI scientists and change leaders.
  • Opportunity to work on research questions with long shelf-life and wide applicability.
  • Aim for publications at competitive venues.
  • Culture of research excellence.
  • Hybrid setup in Gothenburg or a remote setup from Stockholm.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service