AION
Open-source releaseEfficiency & Inference · MLOps, Tooling & Infrastructure1 source · May 29, 2026

vllm-project/vllm v0.22.0

DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).

Key points

  • This release features 459 commits from 230 contributors (63 new)!
  • A large set of fused kernels (MegaMoE, mhc, Q-norm, indexer, sparse MLA) and ROCm parity fixes landed alongside accuracy fixes (#42810, #43710).
  • Model Runner V2 advances toward default: MRv2 is now default for Qwen3 dense models. vLLM will fall back to MRv1 for features that aren't yet supported in MRv2 (#39337). sleep-mode weight reload (#42673), updateconfig (#42783), and shared KV-cache layers (#35045), plus many correctness fixes.
  • Experimental Rust frontend: A new Rust front-end integration landed (#40848), with the implementation moved into the tree (#43283) and a DP Supervisor for data-parallel serving (#40841).

Sources (1)

  • [1]vllm-project/vllm v0.22.0
    GitHub: vllm-project/vllm · May 29, 10:28 AM
    * **DeepSeek V4 maturity**: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated `vllm/models/deepseek_v4/` package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
    This release features 459 commits from 230 contributors (63 new)!

Extractive summary: sentences quoted from the sources.

Before this

  1. May 16, 2026sgl-project/sglang v0.5.12
  2. May 15, 2026vllm-project/vllm v0.21.0
  3. May 5, 2026sgl-project/sglang v0.5.11
  4. Apr 27, 2026vllm-project/vllm v0.20.0
  5. Apr 3, 2026vllm-project/vllm v0.19.0
  6. Feb 24, 2026sgl-project/sglang v0.5.9

Related