AION
Open-source releaseLarge Language Models1 source · Sep 18, 2026

sgl-project/sglang v0.5.20

| Model | Type | PRs | Cookbook |

Key points

  • Masks now run under overlap scheduling: on Qwen3-8B, decode throughput is 17% higher at batch 1 and 52% higher at batch 64 than the previous implementation.
  • DSpark under PD with decode context parallelism. A DCP1 prefill can now transfer its DSpark draft KV to a DCP-N decode, so hybrid models such as Kimi-Linear run DSpark in disaggregated, context-parallel serving.
  • Responses API storage is opt-in. /v1/responses no longer retains results in memory unless the server starts with --enable-response-store.
  • SGLang Simulator. A CPU-only simulator runs the real scheduler, radix cache, and hierarchical cache with a latency predictor in place of the model forward.

Sources (1)

  • [1]sgl-project/sglang v0.5.20
    GitHub: sgl-project/sglang · Sep 18, 10:41 PM
    | Model | Type | PRs | Cookbook |
    Masks now run under overlap scheduling: on Qwen3-8B, decode throughput is 17% higher at batch 1 and 52% higher at batch 64 than the previous implementation.

Extractive summary: sentences quoted from the sources.