AION
Open-source releaseEfficiency & Inference1 source · Oct 5, 2026

vllm-project/vllm v0.31.0

Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).

Key points

  • This release features 717 commits from 307 contributors (96 new)!
  • Experimental initialized-engine snapshots (vllm snapshot create/restore) use CRIU to restore a fully initialized TP1 engine (#51360).
  • Breaking changes: per-request multimodal kwargs gated (#58830); tokenizermode="slow" removed (#58545); --enable-mamba-fine-grained-prefix-cache renamed to --enable-mamba-shared-prefix-checkpoint (#57382); online quantization through quantization="fp8" replaced by the fp8pertensor shorthand (#53585) and Quark silent online quantization removed (#51800); the AllSpark INT8 W8A16 backend removed (#58001); --enforce-eager now also disables JIT kernel warmup (#58197); XPU graphs enabled by default with VLLMXPUENABLEXPUGRAPH removed (#51600).
  • | CUDA 13.0 (Default) | docker pull vllm/vllm-openai:v0.31.0 |

Sources (1)

  • [1]vllm-project/vllm v0.31.0
    GitHub: vllm-project/vllm · Oct 5, 06:44 AM
    * **Fast restart**: the new `vllm preload` CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a `/health` endpoint (#58552) and a readiness wait (#58370).
    This release features 717 commits from 307 contributors (96 new)!

Extractive summary: sentences quoted from the sources.