Repository / library

SGLang

3stories this week
6last 30 days
25all time

Timeline

  1. Oct 7, 2026 · Research paper · 1 source
    Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches
    We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.
  2. Oct 6, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.903-beta: New Browser + Voice Cloning
    This release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
  3. Oct 6, 2026 · Research paper · 1 source
    VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
    Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.
  4. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  5. Sep 28, 2026 · Opinion / analysis · 1 source
    How GLM5.3 Sparse Attention Affects HBM Memory Usage
    How Sparse Attention Affects DRAM/NAND Memory
  6. Sep 18, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.20
    | Model | Type | PRs | Cookbook |
  7. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  8. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  9. Aug 8, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.17
    Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
  10. Jul 25, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.16
    DSpark: confidence-driven speculative decoding: A new speculative algorithm.
  11. Jul 14, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15.post1
    v0.5.15.post1 includes a few patches, mostly for GLM 5.2
  12. Jul 10, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15
    GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
  13. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  14. Jun 13, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.13
    DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
  15. Jun 9, 2026 · Model release · 1 source
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model
  16. May 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12.post1
    v0.5.12.post1 is a stability patch on top of v0.5.12.
  17. May 16, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12
    DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
  18. May 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.11
    Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
  19. Apr 9, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10.post1
    Bumps flashinfer from v0.6.7.post2 to v0.6.7.post3 to resolve an issue in its jit cubin downloader.
  20. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  21. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  22. Feb 24, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.9
    TRT-LLM NSA Kernel Integration for DeepSeek V3.2: Integrate TRT-LLM DSA kernels for Native Sparse Attention, boosting DeepSeek V3.2 performance by 3x-5x on Blackwell platforms with trtllm for both --nsa-prefill-backend and --nsa-decode-backend
  23. Jan 23, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.8
    Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
  24. Jan 9, 2026 · Open-source release · 1 source
    sgl-project/sglang gateway-v0.3.1: Release Gateway-v0.3.1
    We're excited to announce SMG v0.3.1 – a game-changing release with 10-12x performance improvement and 99% memory reduction in cache-aware routing, plus enterprise-grade security!
  25. Jan 1, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.7
    [SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered

Often appears with