Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning.
ProofPaper ↗
Key points
- Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility.
- Nullify employs steering vectors during inference to redirect privacy-related activations away from their memorized answers, while satisfying a null-space constraint that leaves retained-query activations essentially unaffected to maintain utility.
- Evaluations on TOFU and MUSE show that Nullify matches or surpasses established baselines in forget quality while achieving near-lossless preservation of model utility.
- By avoiding weight updates entirely, Nullify serves as an efficient, plug-and-play inference-time intervention framework.
Sources (1)
- [1]Nullify: Null-Space Activation Steering for Training-Free LLM UnlearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:24 PM
To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning.
Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching
- Oct 6, 2026Unlocking Earth AI’s planetary geospatial foundation models for global public health
- Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- Oct 1, 2026Introducing Clef: our open-source decision models, and new RL fine-tuning platform
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK