AION
Open-source releaseEfficiency & Inference1 source · Aug 26, 2026

vllm-project/vllm v0.28.0

DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).

Key points

  • This release features 584 commits from 270 contributors (76 new)!
  • Kimi-K3 also now runs on ROCm with the V2 model runner (#51653).
  • Rust frontend & gRPC: a standalone renderer (#50289), multimodal image inference over gRPC (#50368), explicit data-parallel rank routing (#51178), and RL lifecycle control (#51316), with protobuf schemas now published to Buf (#51276).
  • | CUDA 12.9 | docker pull vllm/vllm-openai:v0.28.0-cu129 |

Sources (1)

  • [1]vllm-project/vllm v0.28.0
    GitHub: vllm-project/vllm · Aug 26, 09:46 AM
    * **DeepSeek V4**: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
    This release features 584 commits from 270 contributors (76 new)!

Extractive summary: sentences quoted from the sources.