A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks
Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data.
Key points
- In this data-free regime, we analyze where forgetting occurs and why.
- Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere.
- Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates.
- We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training.
Sources (1)
- [1]A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never SpeaksarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:01 AM
Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data.
In this data-free regime, we analyze where forgetting occurs and why.
Extractive summary: sentences quoted from the sources.