AION
Open-source releaseTraining & Scaling · Large Language Models · Efficiency & Inference1 source · Jul 10, 2026

sgl-project/sglang v0.5.15

GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.

Key points

  • It now runs at 500+ tok/s/user on 8x B300, 450 on 4x GB300 (bs=1).
  • Full release notes by category below.

Sources (1)

  • [1]sgl-project/sglang v0.5.15
    GitHub: sgl-project/sglang · Jul 10, 10:58 PM
    **GLM-5.2 NVFP4, tuned for production**: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
    It now runs at **500+ tok/s/user on 8x B300, 450 on 4x GB300** (bs=1).

Extractive summary: sentences quoted from the sources.