vllm-project/vllm v0.28.0
DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
Key points
- This release features 584 commits from 270 contributors (76 new)!
- Kimi-K3 also now runs on ROCm with the V2 model runner (#51653).
- Rust frontend & gRPC: a standalone renderer (#50289), multimodal image inference over gRPC (#50368), explicit data-parallel rank routing (#51178), and RL lifecycle control (#51316), with protobuf schemas now published to Buf (#51276).
- | CUDA 12.9 | docker pull vllm/vllm-openai:v0.28.0-cu129 |
Sources (1)
- [1]vllm-project/vllm v0.28.0GitHub: vllm-project/vllm · Aug 26, 09:46 AM
* **DeepSeek V4**: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
This release features 584 commits from 270 contributors (76 new)!
Extractive summary: sentences quoted from the sources.