Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell
Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds.
Key points
- Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain.
- We measure both on one three-tier agent architecture.
- Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines.
- We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve.
Sources (1)
- [1]Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can TellarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:17 AM
Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds.
Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain.
Extractive summary: sentences quoted from the sources.