KV cache
Also known as: KV-cache, key-value cache
19stories this week
23last 30 days
46all time
Timeline
- Oct 10, 2026 · Opinion / analysis · 1 source48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
- Oct 10, 2026 · Opinion / analysis · 1 sourceStrata with Qwen3.8 Flash Next UD-Q4_K_XLMost of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
- Oct 8, 2026 · Research paper · 1 sourceCompile the Table: Query-Calibrated Operator Compression for Tabular In-Context LearningWe propose QCOC (Query-Calibrated Operator Compression), which exploits the exchangeability and repeated use of in-context examples by compiling their full KV cache once into compact memory shared across subsequent queries.
- Oct 8, 2026 · Research paper · 1 sourceMemory Forcing: Attendable Mid-Horizon History for Streaming Video GenerationAutoregressive video diffusion enables causal video streaming without a bidirectional pass over the full clip, but existing few-step systems usually retain only the opening and most recent frames in a fixed-size KV cache.
- Oct 8, 2026 · Research paper · 1 sourceRaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective RecomputationTo bridge this gap, we introduce RaReCache, a framework that enables a large target model to decode accurately from a cache prefilled by a much smaller source via selective recomputation.
- Oct 8, 2026 · Research paper · 1 sourceRead What Matters: Query-Adaptive Quantization for KV CachesKV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
- Oct 8, 2026 · Research paper · 1 sourceBridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight MigrationKV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference.
- Oct 7, 2026 · Research paper · 1 sourceResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit ResidualsLooped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count.
- Oct 7, 2026 · Research paper · 1 sourceDual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV CachesWe introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.
- Oct 7, 2026 · Research paper · 1 sourceCHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model EngineeringGeometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases.
- Oct 6, 2026 · Research paper · 1 sourceA Self-Pruning Transformer: Extreme KV-Cache Compression with Universal AttentionRecent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.
- Oct 6, 2026 · Research paper · 1 sourceVisual Memory Attacks Can Persist Through The KV CacheWe then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention.
- Oct 6, 2026 · Research paper · 1 sourceSPIN: Shadow Predictive Indexer for Sparse AttentionWe propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead.
- Oct 6, 2026 · Research paper · 1 sourceEnabling Dynamic Computation in Looped LMsLooped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens).
- Oct 6, 2026 · Research paper · 1 sourceHybrid Latent Attention for Looped Language ModelsWe propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values.
- Oct 6, 2026 · Research paper · 1 sourceReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon AgentsTo overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.
- Oct 6, 2026 · Research paper · 1 sourcePersistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can TellDecomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds.
- Oct 6, 2026 · Research paper · 1 sourceTRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language ModelsIn this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods.
- Oct 5, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.31.0Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
- Sep 30, 2026 · Research paper · 1 sourceSparseEngine: Sparse-First Inference EngineWe present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 28, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.900-beta: Laya Decision Models + LibraryRun and serve Decision Models like Laya (open-source Jev) locally
- Sep 22, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.30.0This release features 762 commits from 315 contributors (104 new)!
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 26, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.28.0DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
- Aug 21, 2026 · Open-source release · 1 sourceollama/ollama v0.33.0Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).