Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel.
ProofPaper ↗
Key points
- We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding.
- We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus.
- We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding.
- Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency.
Sources (1)
- [1]Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact DistributionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:18 AM
Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel.
We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 22, 2026vllm-project/vllm v0.30.0
- Sep 9, 2026vllm-project/vllm v0.29.0
- Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
- Aug 26, 2026vllm-project/vllm v0.28.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 29, 2026vllm-project/vllm v0.24.0