AION
Research paperEfficiency & Inference · Large Language Models · Agents & Tool Use1 source · Sep 30, 2026

SparseEngine: Sparse-First Inference Engine

We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.

Key points

  • Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation.
  • Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specific layouts or workflows.
  • SparseEngine supports 15 methods across four categories and enables cross-request state management through Chain Cache, which resumes KV-eviction methods from retained history, and controllable Prefix-Cache Pruning, which removes KV from selected history regions while preserving logical-prefix matching.
  • While maintaining method quality, SparseEngine delivers over 10x higher throughput with KV eviction, over 2.5x faster decoding at matched concurrency than vLLM, and over 2x end-to-end speedup on agent benchmarks.

Sources (1)

  • [1]SparseEngine: Sparse-First Inference Engine
    Hugging Face Daily Papers · Sep 30, 12:00 AM
    We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.
    Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation.

Extractive summary: sentences quoted from the sources.