Open sourceOpen-source releaseEfficiency & Inference1 source · Aug 8, 2026

sgl-project/sglang v0.5.17

Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.

Key points

  • MiniMax-H3 day-0 support: MiniMax's video generation model that produces a video and a synchronized stereo audio track in one request, served natively on SGLang-Diffusion across all three public task profiles: text-to-video-and-audio (t2va), first/last-frame conditioning (fl2va), and image/video/audio reference conditioning (ref2va, which also covers video-to-video).
  • Session-reference-aware Unified Radix Cache: For agentic and RL-rollout workloads, requests can carry a stable sessionid so eviction knows which prefixes an active session still references, instead of evicting purely by cache policy.
  • SM90 FP8 MegaMoE for DeepSeek-V4: Adds the DeepGEMM MegaMoE A2A path on SM90 for DeepSeek-V4-Flash/Pro FP8, including the pre-dispatch JIT kernel and FP8 expert weight preparation.
  • Faster engine recovery: Large-model restarts cost 3 to 6+ minutes today, about 6.5 minutes for Qwen3-235B FP8 on 4 GPUs, because weights reload from storage and CUDA graphs recapture.

Sources (1)

  • [1]sgl-project/sglang v0.5.17
    GitHub: sgl-project/sglang · Aug 8, 12:19 AM
    **Kimi K3 day-0 support**: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
    **MiniMax-H3 day-0 support**: MiniMax's video generation model that produces a video and a synchronized stereo audio track in one request, served natively on SGLang-Diffusion across all three public task profiles: text-to-video-and-audio (`t2va`), first/last-frame conditioning (`fl2va`), and image/video/audio reference conditioning (`ref2va`, which also covers video-to-video).

Extractive summary: sentences quoted from the sources.

Before this

  1. Jul 10, 2026sgl-project/sglang v0.5.15
  2. Jun 29, 2026vllm-project/vllm v0.24.0
  3. Jun 26, 2026sgl-project/sglang v0.5.14
  4. Jun 13, 2026sgl-project/sglang v0.5.13
  5. May 16, 2026sgl-project/sglang v0.5.12
  6. May 5, 2026sgl-project/sglang v0.5.11

Related