ResearchResearch paperEfficiency & Inference1 source · Oct 8, 2026

Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding

Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch.

Key points

  • We ask whether better exact verification can increase acceptance on a fixed tree and how target feedback can improve the tree itself.
  • Through a target-flow view, we identify a canonical exit law and prove that one plus target coverage sharply bounds the expected output-block length, including the bonus token, of any exact path verifier.
  • This yields Tree Exit Verification (TEV), an exact, level-parallel procedure using one exit-node decision and one bonus-token decision.
  • Experiments across dialogue, code, and mathematical reasoning validate fixed-tree equivalence: ExitTrain increases average output-block length by 13%, while TEV reduces verifier-stage latency by 15%, yielding a 14% end-to-end speedup over DDTree.

Sources (1)

  • [1]Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:40 AM
    Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch.
    We ask whether better exact verification can increase acceptance on a fixed tree and how target feedback can improve the tree itself.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Speedbumps: Rejection Attacks on Speculative Decoding
  2. Oct 7, 2026Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
  3. Oct 6, 2026SPIN: Shadow Predictive Indexer for Sparse Attention
  4. Oct 6, 2026Secure Speculative Decoding for Large Language Models

Related