AION
Open-source releaseLarge Language Models1 source · Aug 22, 2026

sgl-project/sglang v0.5.18

| Model | Type | PRs | Cookbook |

Key points

  • FlashInfer MNNVL for pure allreduce: Non-fused allreduce sites now reuse the FlashInfer MNNVL workspace instead of falling back to NCCL.
  • DeepSeek-V4-Flash TP4 decode on Blackwell gains up to +6.9% at small batches.
  • NVFP4 checkpoints run on AMD: --quantization quarkmxfp4 dequantizes ModelOpt and Quark NVFP4 weights and requantizes to MXFP4 at load, never holding a full-precision copy.
  • Kimi K3 tuned for MI355X: a grouped-head MLA verify kernel replaces the MHA-shaped split-KV path that re-read the shared latent once per head, for 1.37-1.77x throughput and 1.45-2.42x ITL at concurrency 2-32; the AITER MLA prefill kernel now accepts K3's 12-head shape (TTFT up to -14.9%); a gfx950-tuned decode stage-1 geometry adds 46-73% ITL at 68k input, opt-in via SGLANGMLADECODETUNE=1.

Sources (1)

  • [1]sgl-project/sglang v0.5.18
    GitHub: sgl-project/sglang · Aug 22, 12:09 AM
    | Model | Type | PRs | Cookbook |
    **FlashInfer MNNVL for pure allreduce**: Non-fused allreduce sites now reuse the FlashInfer MNNVL workspace instead of falling back to NCCL.

Extractive summary: sentences quoted from the sources.