sgl-project/sglang v0.5.16
DSpark: confidence-driven speculative decoding: A new speculative algorithm.
Key points
- It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft length.
- Reaches 383.7 tok/s at accept length ~5 on DeepSeek-V4-Pro, TP8 on B300 (bs=1).
- GLM-5.2 DSA cache layer split under prefill CP: KV and indexer cache layers are sharded across CP ranks.
- ReplaySSM Ring Spec-Verify (GDN): Drops the per-draft SSM snapshot.
Sources (1)
- [1]sgl-project/sglang v0.5.16GitHub: sgl-project/sglang · Jul 25, 12:13 AM
**DSpark: confidence-driven speculative decoding**: A new speculative algorithm.
It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft length.
Extractive summary: sentences quoted from the sources.