sgl-project/sglang v0.5.8
Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
Key points
- Day 0 Support for GLM 4.7 Flash: #17247
- Context Parallelism Optimization with support for fused MoE, multi-batch, and FP8 KV cache: #13959
- Added dependencies for tvm-ffi and quack-kernels: #17075
- Mooncake transfer engine updated to 0.3.8.post1: #16792
Sources (1)
- [1]sgl-project/sglang v0.5.8GitHub: sgl-project/sglang · Jan 23, 10:09 PM
* Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
* Day 0 Support for GLM 4.7 Flash: #17247
Extractive summary: sentences quoted from the sources.