vllm-project/vllm v0.25.0
Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
Key points
- This release features 558 commits from 232 contributors (64 new)!
- Model Runner V2 is now the default for all dense models (#44443).
- GLM-5 / DeepSeek-V3.2 landed in the model zoo (#46808) with GLM-5.2 tuning, and MiniMax-M3 gained pipeline parallelism (#45810) and NVFP4 support (#46756).
- Transformers backend: now as fast as native vLLM (#47187), FP8 MoE fix (#46820), embed scaling + CUDA graph fix (#48010), GPTBigCode/Starcoder2 (#30966) and RoBERTa (#47452) migration, M-RoPE mmtokentypeids fix (#46552), tied-embedding lmhead.bias fix (#46835).
Sources (1)
- [1]vllm-project/vllm v0.25.0GitHub: vllm-project/vllm · Jul 11, 08:06 PM
Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
This release features 558 commits from 232 contributors (64 new)!
Extractive summary: sentences quoted from the sources.