SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
Key points
- Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding.
- To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding.
- At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal.
- Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.
Sources (2)
- [1]SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferenceHugging Face Daily Papers · Oct 8, 12:00 AM
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding.
- [2]SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferencearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:06 PM · same content
Extractive summary: sentences quoted from the sources.