AION
Open-source releaseEfficiency & Inference · MLOps, Tooling & Infrastructure1 source · Jun 5, 2026

vllm-project/vllm v0.22.1

v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions.

Key points

  • This release features 8 commits from 6 contributors (1 new)!
  • New model: JetBrains' Mellum v2, an open-weights Mixture-of-Experts code-generation model (#43992).
  • Fix HyperCLOVAX loading after the upstream HuggingFace repo removed its remote code (now native in transformers >= 5.9.0): register the hyperclovax modeltype so vLLM uses its vendored config instead of the stale automap (#43860).
  • Fix a deterministic hang in multi-node Ray data-parallel serving with numapiservers > 1 by excluding the Ray DP backend from the deferred (kernel-assigned) port allocation introduced in #42585 (#43864).

Sources (1)

  • [1]vllm-project/vllm v0.22.1
    GitHub: vllm-project/vllm · Jun 5, 10:10 AM
    v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions.
    This release features 8 commits from 6 contributors (1 new)!

Extractive summary: sentences quoted from the sources.