SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning
We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B.
Key points
- We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference.
- On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws.
- An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone.
- We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.
Sources (1)
- [1]SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and ReasoningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:00 PM
We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B.
We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference.
Extractive summary: sentences quoted from the sources.