sgl-project/sglang v0.5.18
| Model | Type | PRs | Cookbook |
Key points
- FlashInfer MNNVL for pure allreduce: Non-fused allreduce sites now reuse the FlashInfer MNNVL workspace instead of falling back to NCCL.
- DeepSeek-V4-Flash TP4 decode on Blackwell gains up to +6.9% at small batches.
- NVFP4 checkpoints run on AMD: --quantization quarkmxfp4 dequantizes ModelOpt and Quark NVFP4 weights and requantizes to MXFP4 at load, never holding a full-precision copy.
- Kimi K3 tuned for MI355X: a grouped-head MLA verify kernel replaces the MHA-shaped split-KV path that re-read the shared latent once per head, for 1.37-1.77x throughput and 1.45-2.42x ITL at concurrency 2-32; the AITER MLA prefill kernel now accepts K3's 12-head shape (TTFT up to -14.9%); a gfx950-tuned decode stage-1 geometry adds 46-73% ITL at 68k input, opt-in via SGLANGMLADECODETUNE=1.
Sources (1)
- [1]sgl-project/sglang v0.5.18GitHub: sgl-project/sglang · Aug 22, 12:09 AM
| Model | Type | PRs | Cookbook |
**FlashInfer MNNVL for pure allreduce**: Non-fused allreduce sites now reuse the FlashInfer MNNVL workspace instead of falling back to NCCL.
Extractive summary: sentences quoted from the sources.