vllm-project/vllm v0.20.0
CUDA 13.0 default: Default CUDA wheel on PyPI and vllm/vllm-openai:v0.20.0 image switched to CUDA 13.0; architecture lists and build-args cleaned up (#39878), and CUDA bumped to 13.0.2 to match PyTorch 2.11.0 (#40669).
Key points
- This release features 752 commits from 320 contributors (123 new)!
- We highly recommend to install vLLM with uv and use --torch-backend=cu129 if you are on CUDA 12.9.
- Transformers v5: vLLM now runs on HuggingFace transformers>=5 (#30566), with vision-encoder torch.compile bypass (#30518) and continued v4/v5 compat fixes including PaddleOCR-VL image processor maxpixels (#38629), Mistral YaRN warning (#37292), and Jina ColBERT rotary invfreq recompute (#39176).
- New architectures: DeepSeek V4 (#40860), Hunyuan v3 preview (#40681), Granite 4.1 Vision (#40282), EXAONE-4.5 (#39388), BharatGen Param2MoE (#38000), Phi-4-reasoning-vision-15B (#38306), Cheers multimodal (#38788), telechat3 (#38510), FireRedLID (#39290), jina-reranker-v3 (#38800), Jina Embeddings v5 (#39575), Nemotron-v3 VL Nano/Super (#39747).
Sources (1)
- [1]vllm-project/vllm v0.20.0GitHub: vllm-project/vllm · Apr 27, 09:20 PM
* **CUDA 13.0 default**: Default CUDA wheel on PyPI and `vllm/vllm-openai:v0.20.0` image switched to CUDA 13.0; architecture lists and build-args cleaned up (#39878), and CUDA bumped to 13.0.2 to match PyTorch 2.11.0 (#40669).
This release features 752 commits from 320 contributors (123 new)!
Extractive summary: sentences quoted from the sources.