RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
To bridge this gap, we introduce RaReCache, a framework that enables a large target model to decode accurately from a cache prefilled by a much smaller source via selective recomputation.
Key points
- Cross-model KV-cache reuse remains a key challenge in modern LLM serving.
- Coding agents and multi-model systems increasingly route a shared context across models: a user may switch models mid-session, or a cascade may escalate a difficult query.
- Recent work shows that closed-form linear maps can translate KV caches between models in the same family, but transfer accuracy degrades as the model-size gap widens.
- RaReCache establishes an efficient serving paradigm where small models prefill on behalf of massive targets, enabling large models to recompute only critical tokens, drastically reducing prefill latency.
Sources (1)
- [1]RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective RecomputationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:49 AM
To bridge this gap, we introduce RaReCache, a framework that enables a large target model to decode accurately from a cache prefilled by a much smaller source via selective recomputation.
Cross-model KV-cache reuse remains a key challenge in modern LLM serving.
Extractive summary: sentences quoted from the sources.