AION
Research paperLarge Language Models · Efficiency & Inference1 source · Oct 7, 2026

OrBIT: Structure-Guided Embedding Compression

We introduce OrBIT, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords.

Key points

  • Embedding tables are among the largest components of modern language models.
  • Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it.
  • Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec.
  • Across four LLM embedding tables, OrBIT achieves $37.9\times$ compression on GPT-2 and over $23\times$ on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.

Sources (1)

  • [1]OrBIT: Structure-Guided Embedding Compression
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:44 PM
    We introduce OrBIT, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords.
    Embedding tables are among the largest components of modern language models.

Extractive summary: sentences quoted from the sources.