vllm-project/vllm v0.22.0
DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
Key points
- This release features 459 commits from 230 contributors (63 new)!
- A large set of fused kernels (MegaMoE, mhc, Q-norm, indexer, sparse MLA) and ROCm parity fixes landed alongside accuracy fixes (#42810, #43710).
- Model Runner V2 advances toward default: MRv2 is now default for Qwen3 dense models. vLLM will fall back to MRv1 for features that aren't yet supported in MRv2 (#39337). sleep-mode weight reload (#42673), updateconfig (#42783), and shared KV-cache layers (#35045), plus many correctness fixes.
- Experimental Rust frontend: A new Rust front-end integration landed (#40848), with the implementation moved into the tree (#43283) and a DP Supervisor for data-parallel serving (#40841).
Sources (1)
- [1]vllm-project/vllm v0.22.0GitHub: vllm-project/vllm · May 29, 10:28 AM
* **DeepSeek V4 maturity**: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated `vllm/models/deepseek_v4/` package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
This release features 459 commits from 230 contributors (63 new)!
Extractive summary: sentences quoted from the sources.
Before this
- May 16, 2026sgl-project/sglang v0.5.12
- May 15, 2026vllm-project/vllm v0.21.0
- May 5, 2026sgl-project/sglang v0.5.11
- Apr 27, 2026vllm-project/vllm v0.20.0
- Apr 3, 2026vllm-project/vllm v0.19.0
- Feb 24, 2026sgl-project/sglang v0.5.9