GLM
10stories this week
15last 30 days
34all time
In the model registry
| Model | Params | Context | Released |
|---|---|---|---|
| Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw | 146B | 1M | Oct 1, 2026 |
| nvidia/GLM-5.3-NVFP4 | 391B | 1M | Sep 14, 2026 |
| nvidia/GLM-5.3-Flash-NVFP4 | 169B | 1M | Sep 2, 2026 |
| zai-org/GLM-5.3-Flash-BF16 | 321B | 1M | Aug 25, 2026 |
| zai-org/GLM-5.3-BF16 | 753B | 1M | Aug 25, 2026 |
| zai-org/GLM-5.3-Flash | 321B | 1M | Aug 25, 2026 |
| zai-org/GLM-5.3 | 753B | 1M | Aug 25, 2026 |
| zai-org/GLM-5.2 | 753B | 1M | Jun 16, 2026 |
| zai-org/GLM-5.2-FP8 | 753B | 1M | Jun 16, 2026 |
| zai-org/GLM-5.1-FP8 | 754B | 198K | Apr 3, 2026 |
| zai-org/GLM-5.1 | 754B | 198K | Apr 3, 2026 |
| zai-org/GLM-5 | 754B | 198K | Feb 11, 2026 |
| zai-org/GLM-5-FP8 | 754B | 198K | Feb 11, 2026 |
| zai-org/GLM-OCR | 1.3B | 128K | Jan 30, 2026 |
| zai-org/GLM-4.7-Flash | 31.2B | 198K | Jan 19, 2026 |
| zai-org/GLM-Image | 6.9B | – | Jan 8, 2026 |
| zai-org/GLM-4.7-FP8 | 358B | 198K | Dec 22, 2025 |
| zai-org/GLM-4.7 | 358B | 198K | Dec 22, 2025 |
| zai-org/GLM-TTS | – | – | Dec 10, 2025 |
| zai-org/GLM-ASR-Nano-2512 | 2.3B | 8K | Dec 9, 2025 |
| zai-org/GLM-4.6V-FP8 | 108B | 128K | Dec 7, 2025 |
| zai-org/GLM-4.6V | 108B | 128K | Dec 7, 2025 |
| zai-org/GLM-4.6V-Flash | 10.3B | 128K | Dec 7, 2025 |
| zai-org/GLM-4.6 | 357B | 198K | Sep 29, 2025 |
Timeline
- Oct 8, 2026 · Research paper · 1 sourceSafe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard ModelsIn this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations.
- Oct 8, 2026 · Research paper · 2 sourcesREMORY: Learning Residual Memory for Context CompactionWe introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
- Oct 8, 2026 · Research paper · 1 sourceQUILT: Rethinking Sparse-Attention Prefill through Shared Query ExecutionWe present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
- Oct 7, 2026 · Research paper · 1 sourceHow Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and DefenseSafety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.
- Oct 6, 2026 · Research paper · 1 sourceSocio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability DistillationSimulating individual behavior requires large language models (LLMs) to preserve persona traits while adapting to dynamic social contexts.
- Oct 6, 2026 · Research paper · 1 sourceMINDSET: Energy-based Schema Evolution for Long Conversational Agent MemoryWe introduce MINDSET, a memory controller that stores a conversation as immutable episodes and organizes them into versioned schemas through minimum-energy state transitions.
- Oct 6, 2026 · Opinion / analysis · 1 sourceThe Cyber Risk Discourse is BrokenWestern voices saying open weight models are necessary for defense and banning them will make the world less safe, occupied by AI risk moderates to different extremes.
- Oct 6, 2026 · Research paper · 1 sourceImage Bitstream Fine-grained Understanding for Privacy-Friendly AIoTImage Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences.
- Oct 6, 2026 · Research paper · 1 sourceServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks.
- Oct 5, 2026 · Product / feature launch · 1 sourceIntroducing GLM 5.3 on Amazon BedrockGLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock.
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Oct 1, 2026 · Open-source release · 1 sourceInfatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpwInfatoshi published the model GLM-5.3-UNCENSORED-EXL3-3.0bpw on Hugging Face.
- Sep 30, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.18.0: Release 5.18.0Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
- Sep 22, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.30.0This release features 762 commits from 315 contributors (104 new)!
- Sep 18, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.20| Model | Type | PRs | Cookbook |
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 26, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.16.1: Release v5.16.1This is a special release as we include GLM! (and a few small fixes)
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jul 25, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.16DSpark: confidence-driven speculative decoding: A new speculative algorithm.
- Jul 14, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15.post1v0.5.15.post1 includes a few patches, mostly for GLM 5.2
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jul 10, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.23.0DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
- May 16, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- May 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.11Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
- Apr 6, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331