Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning
Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities.
ProofPaper ↗
Key points
- Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior.
- Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance.
- In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities.
- PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints.
Sources (1)
- [1]Not Every Change Is Necessary: Recoverable Drift in Large Language Model UnlearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:12 PM
Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities.
Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior.
Extractive summary: sentences quoted from the sources.