AION
Open-source releaseMLOps, Tooling & Infrastructure · Efficiency & Inference1 source · Aug 10, 2026

vllm-project/vllm v0.27.0

Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).

Key points

  • This release features 561 commits from 242 contributors (64 new)!
  • More new models: Qwen3.5 text-only dense and MoE models (#50210) with EVS video token pruning (#48912), K-EXAONE-2.0-750B-A37B (#50524), VaultGemma via the Transformers modeling backend (#49803), and jina-embeddings-v5-text-nano (#50688).
  • Rust frontend grows a gRPC control plane: engine-aware health reporting (#48992), abort control (#49255), server and model discovery (#49491), KV event source discovery (#50033), plus vllm-bench integrated into the vllm CLI (#48930).
  • Inkling: llm-compressor NVFP4 weights (#49258) and compressed-tensors dynamic FP8 (#48876).

Sources (1)

  • [1]vllm-project/vllm v0.27.0
    GitHub: vllm-project/vllm · Aug 10, 09:18 PM
    * **Kimi K3 support** with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
    This release features 561 commits from 242 contributors (64 new)!

Extractive summary: sentences quoted from the sources.