AION
Research paperLarge Language Models1 source · Oct 7, 2026

SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning

We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B.

Key points

  • We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference.
  • On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws.
  • An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone.
  • We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.

Sources (1)

  • [1]SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:00 PM
    We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B.
    We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference.

Extractive summary: sentences quoted from the sources.