AION
Open-source releaseEfficiency & Inference · MLOps, Tooling & Infrastructure · Large Language Models1 source · Jul 11, 2026

vllm-project/vllm v0.25.0

Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).

Key points

  • This release features 558 commits from 232 contributors (64 new)!
  • Model Runner V2 is now the default for all dense models (#44443).
  • GLM-5 / DeepSeek-V3.2 landed in the model zoo (#46808) with GLM-5.2 tuning, and MiniMax-M3 gained pipeline parallelism (#45810) and NVFP4 support (#46756).
  • Transformers backend: now as fast as native vLLM (#47187), FP8 MoE fix (#46820), embed scaling + CUDA graph fix (#48010), GPTBigCode/Starcoder2 (#30966) and RoBERTa (#47452) migration, M-RoPE mmtokentypeids fix (#46552), tied-embedding lmhead.bias fix (#46835).

Sources (1)

  • [1]vllm-project/vllm v0.25.0
    GitHub: vllm-project/vllm · Jul 11, 08:06 PM
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
    This release features 558 commits from 232 contributors (64 new)!

Extractive summary: sentences quoted from the sources.