vllm-project/vllm v0.29.0
MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
Key points
- This release features 594 commits from 277 contributors (91 new)!
- Model Runner V2 is now the default for all models (#53183), completing the rollout that began with pooling models (#48290).
- New models: Hy4-preview, Tencent's 770B/49B-active MoE with Gated DeepSeek Sparse Attention and native MTP (#54160); Qwen3.8-Flash-Next with BF16/FP8/NVFP4 and MTP (#53896); GraniteSWA and GraniteMoeSWA (#52706); NemotronHOmniReasoningV3 with MTP (#52929, #53121); Kimi K3 NVFP4 checkpoints (#53132).
- Breaking changes: ten deprecated model architectures removed (#53608); FlexOlmo, Olmo3 and Hunyuan V1/VL migrated to the Transformers modeling backend (#53615); PyAV video decoder backend removed (#54231); python -m vllm.entrypoints.openai.apiserver deprecated in favor of vllm serve (#52131); VLLMTESTFORCEFP8MARLIN (#52182) and VLLMROCMUSEAITERFP4ASMGEMM (#53141) removed.
Sources (1)
- [1]vllm-project/vllm v0.29.0GitHub: vllm-project/vllm · Sep 9, 08:54 AM
MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), `extract_hidden_states` speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
This release features 594 commits from 277 contributors (91 new)!
Extractive summary: sentences quoted from the sources.