Skip to main navigation Skip to search Skip to main content

Intrinsic Test of Unlearning Using Parametric Knowledge Traces

  • Yihuai Hong
  • , Lei Yu
  • , Haiqin Yang
  • , Shauli Ravfogel
  • , Mor Geva
  • New York University
  • University of Toronto
  • Shenzhen Technology University
  • Tel Aviv University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The task of "unlearning" certain concepts in large language models (LLMs) has gained attention for its role in mitigating harmful, private, or incorrect outputs. Current evaluations mostly rely on behavioral tests, without monitoring residual knowledge in model parameters, which can be adversarially exploited to recover erased information. We argue that unlearning should also be assessed internally by tracking changes in the parametric traces of unlearned concepts. To this end, we propose a general evaluation methodology that uses vocabulary projections to inspect concepts encoded in model parameters. We apply this approach to localize "concept vectors" - parameter vectors encoding concrete concepts - and construct CONCEPTVECTORS, a benchmark of hundreds of such concepts and their parametric traces in two open-source LLMs. Evaluation on CONCEPTVECTORS shows that existing methods minimally alter concept vectors, mostly suppressing them at inference time, while direct ablation of these vectors removes the associated knowledge and reduces adversarial susceptibility. Our findings reveal limitations of behavior-only evaluations and advocate for parameter-based assessments. We release our code and benchmark at https://github.com/yihuaihong/ConceptVectors.

Original languageEnglish
Title of host publicationEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
EditorsChristos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
PublisherAssociation for Computational Linguistics (ACL)
Pages19513-19535
Number of pages23
ISBN (Electronic)9798891763326
DOIs
StatePublished - 2025
Event30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - Suzhou, China
Duration: 4 Nov 20259 Nov 2025

Publication series

NameEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference

Conference

Conference30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Country/TerritoryChina
CitySuzhou
Period4/11/259/11/25

Bibliographical note

Publisher Copyright:
© 2025 Association for Computational Linguistics.

Fingerprint

Dive into the research topics of 'Intrinsic Test of Unlearning Using Parametric Knowledge Traces'. Together they form a unique fingerprint.

Cite this