AION
Open-source releaseEfficiency & Inference1 source · May 16, 2026

sgl-project/sglang v0.5.12

DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:

Key points

  • DeepGemm and FlashMLA kernels for DeepSeek V4, including MegaMoE
  • Pipeline Parallelism + PD support for DeepSeek-V4: #24700
  • TokenSpeed MLA attention backend (Blackwell, FP8 KV cache): New MLA prefill/decode kernels integrated as an attention backend on SM100, with FP8 KV cache support for low-latency MLA serving: #24925
  • DSv3.2 / GLM-5 FP4 low-latency perf: PDL enabled across DSv3.2 / GLM-5 kernels, torch.mm for the DeepSeek V3.2 indexer GEMM, and a reland of the Cute-DSL FP4 dense GEMM — materially trimming low-latency overheads on FP4 paths: #23965, #23856, #23590, #25311

Sources (1)

  • [1]sgl-project/sglang v0.5.12
    GitHub: sgl-project/sglang · May 16, 06:23 PM
    - **DeepSeek V4 support**: Full inference path for DeepSeek-V4 (#23882), including:
    - DeepGemm and FlashMLA kernels for DeepSeek V4, including MegaMoE

Extractive summary: sentences quoted from the sources.