vllm-project/vllm v0.31.0
Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
Key points
- This release features 717 commits from 307 contributors (96 new)!
- Experimental initialized-engine snapshots (vllm snapshot create/restore) use CRIU to restore a fully initialized TP1 engine (#51360).
- Breaking changes: per-request multimodal kwargs gated (#58830); tokenizermode="slow" removed (#58545); --enable-mamba-fine-grained-prefix-cache renamed to --enable-mamba-shared-prefix-checkpoint (#57382); online quantization through quantization="fp8" replaced by the fp8pertensor shorthand (#53585) and Quark silent online quantization removed (#51800); the AllSpark INT8 W8A16 backend removed (#58001); --enforce-eager now also disables JIT kernel warmup (#58197); XPU graphs enabled by default with VLLMXPUENABLEXPUGRAPH removed (#51600).
- | CUDA 13.0 (Default) | docker pull vllm/vllm-openai:v0.31.0 |
Sources (1)
- [1]vllm-project/vllm v0.31.0GitHub: vllm-project/vllm · Oct 5, 06:44 AM
* **Fast restart**: the new `vllm preload` CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a `/health` endpoint (#58552) and a readiness wait (#58370).
This release features 717 commits from 307 contributors (96 new)!
Extractive summary: sentences quoted from the sources.