AION
Model

Kimi

8stories this week
13last 30 days
32all time

In the model registry

ModelParamsContextReleased
nvidia/Kimi-K3-NVFP41419B1MAug 13, 2026
moonshotai/Kimi-K32780B1MJun 13, 2026
moonshotai/Kimi-K2.7-Code1027B256KJun 11, 2026
moonshotai/Kimi-K2.61027B256KApr 14, 2026
moonshotai/Kimi-K2.51027B256KJan 1, 2026
moonshotai/Kimi-K2-Thinking1026B256KNov 4, 2025
moonshotai/Kimi-Linear-48B-A3B-Base49.1B–Oct 30, 2025
moonshotai/Kimi-Linear-48B-A3B-Instruct49.1B–Oct 30, 2025
moonshotai/Kimi-K2-Instruct-09051026B256KSep 3, 2025
moonshotai/Kimi-K2-Instruct1026B128KJul 11, 2025
moonshotai/Kimi-K2-Base1026B128KJul 3, 2025
moonshotai/Kimi-VL-A3B-Thinking-250616.4B128KJun 21, 2025
moonshotai/Kimi-Dev-72B72.7B128KJun 16, 2025
moonshotai/Kimi-Audio-7B9.8B8KApr 25, 2025
moonshotai/Kimi-Audio-7B-Instruct9.8B8KApr 25, 2025
moonshotai/Kimi-VL-A3B-Thinking16.4B128KApr 9, 2025
moonshotai/Kimi-VL-A3B-Instruct16.4B128KApr 9, 2025

Timeline

  1. Oct 7, 2026 · Research paper · 1 source
    From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs
    To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation.
  2. Oct 7, 2026 · Research paper · 1 source
    Training Advisors for LLM Agents from Task Outcomes
    We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task.
  3. Oct 7, 2026 · Research paper · 1 source
    WorldBench: Evaluating LLMs on Three.js Voxel World Generation
    We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
  4. Oct 7, 2026 · Research paper · 1 source
    Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
    Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
  5. Oct 6, 2026 · Research paper · 1 source
    One Step at a Time: Trading LLM Autonomy for Process Predictability
    Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step.
  6. Oct 6, 2026 · Research paper · 1 source
    ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
    To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks.
  7. Oct 6, 2026 · Research paper · 1 source
    Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
    Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
  8. Oct 5, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.31.0
    Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
  9. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  10. Oct 1, 2026 · Product / feature launch · 1 source
    unslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UX
    This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
  11. Sep 29, 2026 · Open-source release · 1 source
    NVIDIA/TensorRT-LLM v1.3.0rc29
    Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
  12. Sep 22, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.30.0
    This release features 762 commits from 315 contributors (104 new)!
  13. Sep 18, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.20
    | Model | Type | PRs | Cookbook |
  14. Sep 9, 2026 · Open-source release · 1 source
    huggingface/transformers v5.17.0: Release 5.17.0
    Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
  15. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  16. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  17. Aug 26, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.28.0
    DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
  18. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  19. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  20. Aug 8, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.17
    Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
  21. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  22. Jul 10, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15
    GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
  23. Jul 3, 2026 · Open-source release · 1 source
    huggingface/transformers v5.13.0: Release v5.13.0
    Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
  24. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  25. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  26. Jun 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.23.0
    DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
  27. Jun 13, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.13
    DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
  28. May 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.22.0
    DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
  29. May 16, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12
    DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
  30. May 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.21.0
    Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).

Often appears with