AION
Research paperEfficiency & Inference1 source · Oct 7, 2026

Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.

Key points

  • Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache.
  • Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads.
  • Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning.
  • At a 128K context, Dual-QK provides $6.8\times$ KV-cache compression and an estimated $8.3\times$ reduction in KV read volume relative to unpruned BF16.

Sources (1)

  • [1]Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 10:50 AM
    We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.
    Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache.

Extractive summary: sentences quoted from the sources.