AION
Research paperEfficiency & Inference1 source · Oct 8, 2026

Read What Matters: Query-Adaptive Quantization for KV Caches

KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.

Key points

  • We study this mismatch using separate budgets for retained bits and bits fetched per query.
  • ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions.
  • Long-context question answering and retrieval on two instruction-tuned models provide additional quality evidence.
  • On the tested 8K-token, batch-one, single-layer workload on an NVIDIA A10G, a restricted eight-bit ReadKV reader with a two-bit mean payload-read budget has 39% lower latency than the tested TurboQuant codec.

Sources (1)

  • [1]Read What Matters: Query-Adaptive Quantization for KV Caches
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:47 AM
    KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
    We study this mismatch using separate budgets for retained bits and bits fetched per query.

Extractive summary: sentences quoted from the sources.