"Strata" for GLM5.3 Flash is here for some! Project Maya
I stumbled across this as I was currently having glm5.3 flash run only around 10tok/s basically unusable.
Proof1 community thread
Key points
- With Maya i am running 30 tok/s now.
- GLM5.3 feels even with this low quant like a much more enjoyable model than qwen 3.8 flash next so far.
- GLM-5.3-Flash, private and fast · one NVIDIA GPU or up to 16 · AMD and Windows (experimental) · chat, pictures, OpenAI- and Anthropic-compatible API
- Maya runs GLM-5.3-Flash (zai-org, MIT license) on the hardware you already have: 321 B parameters, about 18 B active per token, and a context of up to 1 M tokens.
Sources (1)
- [1]"Strata" for GLM5.3 Flash is here for some! Project Mayar/LocalLLaMA (top, daily) · Oct 11, 06:42 PM
I stumbled across this as I was currently having glm5.3 flash run only around 10tok/s basically unusable.
With Maya i am running 30 tok/s now.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 11, 2026vllm-ascend updates (GLM5.3-flash) on dual 310p Ascend cards
- Oct 11, 2026AMD Reportedly Raises GDDR6 Prices for Board Partners
- Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
- Oct 8, 2026Understanding Jev, the new model everyone is talking about
- Oct 8, 2026The State of AI Report 2026
- Oct 7, 2026[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing