AION
Research paperLarge Language Models1 source · Oct 7, 2026

A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks

Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data.

Key points

  • In this data-free regime, we analyze where forgetting occurs and why.
  • Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere.
  • Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates.
  • We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training.

Sources (1)

Extractive summary: sentences quoted from the sources.