AION
Research paperEfficiency & Inference · Large Language Models1 source · Oct 7, 2026

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.

Key points

  • Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader.
  • A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining.
  • Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank.
  • EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

Sources (1)

  • [1]Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:28 PM
    Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
    Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader.

Extractive summary: sentences quoted from the sources.