vllm-project/vllm v0.19.0
We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
Key points
- This release features 448 commits from 197 contributors (54 new)!
- Zero-bubble async scheduling + speculative decoding: Async scheduling now supports speculative decoding with zero-bubble overlap, significantly improving throughput (#32951).
- Model Runner V2 maturation: MRV2 gains piecewise CUDA graphs for pipeline parallelism (#35162), spec decode rejection sampler with greedy/logprobs support (#37238, #37237), multi-modal embeddings for spec decode (#36097), streaming inputs (#37028), and EPLB support (#37488).
- General CPU KV cache offloading: A simple yet general CPU KV cache offloading mechanism for V1, with pluggable cache policy and block-level preemption handling (#37160, #37874, #34805, #36642, #37853).
Sources (1)
- [1]vllm-project/vllm v0.19.0GitHub: vllm-project/vllm · Apr 3, 02:19 AM
We recommend using pre-built docker image `vllm/vllm-openai:gemma4` for out of box usage.
This release features 448 commits from 197 contributors (54 new)!
Extractive summary: sentences quoted from the sources.