Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
Key points
- Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader.
- A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining.
- Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank.
- EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
Sources (1)
- [1]Cache the Encoder Within:Compact, Reusable Memory across LLM QueriesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:28 PM
Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader.
Extractive summary: sentences quoted from the sources.