AION
Open-source releaseEfficiency & Inference · Large Language Models1 source · Jan 1, 2026

sgl-project/sglang v0.5.7

[SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered

Key points

  • Scalable pipeline parallelism with dynamic chunking support for ultra-long contexts (PP Refactor Roadmap #11857)
  • Encoder Disaggregation for Multi-modal models (Roadmap #15118)
  • Set --dit-layerwise-offload true to reduce peak VRAM usage by up to 30GB, and improve performance by up to 58% for all models
  • Add support for AMD/4090/5090, along with additional attention choices (sage-attn, sage-attn3), more parallelism options (TP) and enhancements to HTTP API (Google vertex supported)

Sources (1)

  • [1]sgl-project/sglang v0.5.7
    GitHub: sgl-project/sglang · Jan 1, 10:01 AM
    - [SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered
    - Scalable pipeline parallelism with dynamic chunking support for ultra-long contexts (PP Refactor Roadmap #11857)

Extractive summary: sentences quoted from the sources.