AION
Model

GLM

10stories this week
15last 30 days
34all time

In the model registry

ModelParamsContextReleased
Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw146B1MOct 1, 2026
nvidia/GLM-5.3-NVFP4391B1MSep 14, 2026
nvidia/GLM-5.3-Flash-NVFP4169B1MSep 2, 2026
zai-org/GLM-5.3-Flash-BF16321B1MAug 25, 2026
zai-org/GLM-5.3-BF16753B1MAug 25, 2026
zai-org/GLM-5.3-Flash321B1MAug 25, 2026
zai-org/GLM-5.3753B1MAug 25, 2026
zai-org/GLM-5.2753B1MJun 16, 2026
zai-org/GLM-5.2-FP8753B1MJun 16, 2026
zai-org/GLM-5.1-FP8754B198KApr 3, 2026
zai-org/GLM-5.1754B198KApr 3, 2026
zai-org/GLM-5754B198KFeb 11, 2026
zai-org/GLM-5-FP8754B198KFeb 11, 2026
zai-org/GLM-OCR1.3B128KJan 30, 2026
zai-org/GLM-4.7-Flash31.2B198KJan 19, 2026
zai-org/GLM-Image6.9B–Jan 8, 2026
zai-org/GLM-4.7-FP8358B198KDec 22, 2025
zai-org/GLM-4.7358B198KDec 22, 2025
zai-org/GLM-TTS––Dec 10, 2025
zai-org/GLM-ASR-Nano-25122.3B8KDec 9, 2025
zai-org/GLM-4.6V-FP8108B128KDec 7, 2025
zai-org/GLM-4.6V108B128KDec 7, 2025
zai-org/GLM-4.6V-Flash10.3B128KDec 7, 2025
zai-org/GLM-4.6357B198KSep 29, 2025

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models
    In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations.
  2. Oct 8, 2026 · Research paper · 2 sources
    REMORY: Learning Residual Memory for Context Compaction
    We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
  3. Oct 8, 2026 · Research paper · 1 source
    QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
    We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
  4. Oct 7, 2026 · Research paper · 1 source
    How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
    Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.
  5. Oct 6, 2026 · Research paper · 1 source
    Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation
    Simulating individual behavior requires large language models (LLMs) to preserve persona traits while adapting to dynamic social contexts.
  6. Oct 6, 2026 · Research paper · 1 source
    MINDSET: Energy-based Schema Evolution for Long Conversational Agent Memory
    We introduce MINDSET, a memory controller that stores a conversation as immutable episodes and organizes them into versioned schemas through minimum-energy state transitions.
  7. Oct 6, 2026 · Opinion / analysis · 1 source
    The Cyber Risk Discourse is Broken
    Western voices saying open weight models are necessary for defense and banning them will make the world less safe, occupied by AI risk moderates to different extremes.
  8. Oct 6, 2026 · Research paper · 1 source
    Image Bitstream Fine-grained Understanding for Privacy-Friendly AIoT
    Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences.
  9. Oct 6, 2026 · Research paper · 1 source
    ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
    To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks.
  10. Oct 5, 2026 · Product / feature launch · 1 source
    Introducing GLM 5.3 on Amazon Bedrock
    GLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock.
  11. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  12. Oct 1, 2026 · Open-source release · 1 source
    Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw
    Infatoshi published the model GLM-5.3-UNCENSORED-EXL3-3.0bpw on Hugging Face.
  13. Sep 30, 2026 · Open-source release · 1 source
    huggingface/transformers v5.18.0: Release 5.18.0
    Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
  14. Sep 22, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.30.0
    This release features 762 commits from 315 contributors (104 new)!
  15. Sep 18, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.20
    | Model | Type | PRs | Cookbook |
  16. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  17. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  18. Aug 26, 2026 · Open-source release · 1 source
    huggingface/transformers v5.16.1: Release v5.16.1
    This is a special release as we include GLM! (and a few small fixes)
  19. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  20. Jul 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.26.0
    New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
  21. Jul 25, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.16
    DSpark: confidence-driven speculative decoding: A new speculative algorithm.
  22. Jul 14, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15.post1
    v0.5.15.post1 includes a few patches, mostly for GLM 5.2
  23. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  24. Jul 10, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15
    GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
  25. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  26. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  27. Jun 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.23.0
    DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
  28. May 16, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12
    DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
  29. May 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.11
    Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
  30. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331

Often appears with