AION
Research paperEfficiency & Inference · Large Language Models2 sources · Oct 8, 2026

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.

Key points

  • Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding.
  • To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding.
  • At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal.
  • Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.

Sources (2)

Extractive summary: sentences quoted from the sources.