Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
Key points
- A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent.
- We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B.
- Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream.
- We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.
Sources (2)
- [1]Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than RecomputeHugging Face Daily Papers · Oct 7, 12:00 AM
We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent.
- [2]Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than RecomputearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:50 PM · same content
Extractive summary: sentences quoted from the sources.