Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding
Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch.
ProofPaper ↗
Key points
- We ask whether better exact verification can increase acceptance on a fixed tree and how target feedback can improve the tree itself.
- Through a target-flow view, we identify a canonical exit law and prove that one plus target coverage sharply bounds the expected output-block length, including the bonus token, of any exact path verifier.
- This yields Tree Exit Verification (TEV), an exact, level-parallel procedure using one exit-node decision and one bonus-token decision.
- Experiments across dialogue, code, and mathematical reasoning validate fixed-tree equivalence: ExitTrain increases average output-block length by 13%, while TEV reduces verifier-stage latency by 15%, yielding a 14% end-to-end speedup over DDTree.
Sources (1)
- [1]Where Draft Trees Lose Target Mass: Exit-Guided Speculative DecodingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:40 AM
Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch.
We ask whether better exact verification can increase acceptance on a fixed tree and how target feedback can improve the tree itself.
Extractive summary: sentences quoted from the sources.