AION
Open-source releaseMLOps, Tooling & Infrastructure · Efficiency & Inference1 source · Sep 9, 2026

vllm-project/vllm v0.29.0

MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).

Key points

  • This release features 594 commits from 277 contributors (91 new)!
  • Model Runner V2 is now the default for all models (#53183), completing the rollout that began with pooling models (#48290).
  • New models: Hy4-preview, Tencent's 770B/49B-active MoE with Gated DeepSeek Sparse Attention and native MTP (#54160); Qwen3.8-Flash-Next with BF16/FP8/NVFP4 and MTP (#53896); GraniteSWA and GraniteMoeSWA (#52706); NemotronHOmniReasoningV3 with MTP (#52929, #53121); Kimi K3 NVFP4 checkpoints (#53132).
  • Breaking changes: ten deprecated model architectures removed (#53608); FlexOlmo, Olmo3 and Hunyuan V1/VL migrated to the Transformers modeling backend (#53615); PyAV video decoder backend removed (#54231); python -m vllm.entrypoints.openai.apiserver deprecated in favor of vllm serve (#52131); VLLMTESTFORCEFP8MARLIN (#52182) and VLLMROCMUSEAITERFP4ASMGEMM (#53141) removed.

Sources (1)

  • [1]vllm-project/vllm v0.29.0
    GitHub: vllm-project/vllm · Sep 9, 08:54 AM
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), `extract_hidden_states` speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
    This release features 594 commits from 277 contributors (91 new)!

Extractive summary: sentences quoted from the sources.