ResearchResearch paperEfficiency & Inference1 source · Oct 6, 2026

Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution

Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel.

Key points

  • We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding.
  • We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus.
  • We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding.
  • Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency.

Sources (1)

  • [1]Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:18 AM
    Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel.
    We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 22, 2026vllm-project/vllm v0.30.0
  2. Sep 9, 2026vllm-project/vllm v0.29.0
  3. Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
  4. Aug 26, 2026vllm-project/vllm v0.28.0
  5. Jul 11, 2026vllm-project/vllm v0.25.0
  6. Jun 29, 2026vllm-project/vllm v0.24.0

Related