ResearchResearch paperEfficiency & Inference1 source · Oct 7, 2026

NeuralZip: Reusable Setup for Fast Lossless Compression

For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios.

Key points

  • Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead.
  • We study whether the statistical structure of exponents can be prepared once and reused.
  • In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction.
  • We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs.

Sources (1)

  • [1]NeuralZip: Reusable Setup for Fast Lossless Compression
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:06 PM
    For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios.
    Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead.

Extractive summary: sentences quoted from the sources.

Related