AION
Open-source releaseEfficiency & Inference1 source · Jul 25, 2026

sgl-project/sglang v0.5.16

DSpark: confidence-driven speculative decoding: A new speculative algorithm.

Key points

  • It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft length.
  • Reaches 383.7 tok/s at accept length ~5 on DeepSeek-V4-Pro, TP8 on B300 (bs=1).
  • GLM-5.2 DSA cache layer split under prefill CP: KV and indexer cache layers are sharded across CP ranks.
  • ReplaySSM Ring Spec-Verify (GDN): Drops the per-draft SSM snapshot.

Sources (1)

  • [1]sgl-project/sglang v0.5.16
    GitHub: sgl-project/sglang · Jul 25, 12:13 AM
    **DSpark: confidence-driven speculative decoding**: A new speculative algorithm.
    It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft length.

Extractive summary: sentences quoted from the sources.