AION
Open-source releaseMLOps, Tooling & Infrastructure · Efficiency & Inference1 source · Sep 22, 2026

vllm-project/vllm v0.30.0

This release features 762 commits from 315 contributors (104 new)!

Key points

  • Watermarking: Gumbel-max watermarked generation and detection with a keyed PRF, per-request opt-out and an example detection endpoint (#54053); dual-key Gumbel-max makes it compatible with speculative decoding (#56122); the Rust frontend forwards the per-request controls (#56338).
  • HiSparse: a host-resident tier for sparse-MLA decode that spills KV pages to pinned host memory under GPU pressure and serves top-k misses from a per-request GPU hot buffer, enabled through HiSparseConnector (#53781), with Prometheus counters (#56061), a host cache shared across TP ranks (#56629), and the attention config inferred from the connector (#57041).
  • Qwen3.8-Flash-Next performance: separate prefill and decode QSA indexer kernels (#54513), fused PLE kernels (#54517), FP8 indexer cache (#54890), padded-index skipping in sparse GQA (#54873), fused PLE residual and QSA output gate (#55309), UVA PLE offload and Engram tensor parallelism via --engram-config (#54371), and torch.compile removed from the NVIDIA implementation so FP8 fits on a single GB300 (#55272).
  • | CUDA 13.0 (Default) | docker pull vllm/vllm-openai:v0.30.0 |

Sources (1)

  • [1]vllm-project/vllm v0.30.0
    GitHub: vllm-project/vllm · Sep 22, 05:20 AM
    This release features 762 commits from 315 contributors (104 new)!
    * **Watermarking**: Gumbel-max watermarked generation and detection with a keyed PRF, per-request opt-out and an example detection endpoint (#54053); dual-key Gumbel-max makes it compatible with speculative decoding (#56122); the Rust frontend forwards the per-request controls (#56338).

Extractive summary: sentences quoted from the sources.