SparseEngine: Sparse-First Inference Engine
We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.
Key points
- Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation.
- Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specific layouts or workflows.
- SparseEngine supports 15 methods across four categories and enables cross-request state management through Chain Cache, which resumes KV-eviction methods from retained history, and controllable Prefix-Cache Pruning, which removes KV from selected history regions while preserving logical-prefix matching.
- While maintaining method quality, SparseEngine delivers over 10x higher throughput with KV eviction, over 2.5x faster decoding at matched concurrency than vLLM, and over 2x end-to-end speedup on agent benchmarks.
Sources (1)
- [1]SparseEngine: Sparse-First Inference EngineHugging Face Daily Papers · Sep 30, 12:00 AM
We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.
Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation.
Extractive summary: sentences quoted from the sources.