Kimi
8stories this week
13last 30 days
32all time
In the model registry
| Model | Params | Context | Released |
|---|---|---|---|
| nvidia/Kimi-K3-NVFP4 | 1419B | 1M | Aug 13, 2026 |
| moonshotai/Kimi-K3 | 2780B | 1M | Jun 13, 2026 |
| moonshotai/Kimi-K2.7-Code | 1027B | 256K | Jun 11, 2026 |
| moonshotai/Kimi-K2.6 | 1027B | 256K | Apr 14, 2026 |
| moonshotai/Kimi-K2.5 | 1027B | 256K | Jan 1, 2026 |
| moonshotai/Kimi-K2-Thinking | 1026B | 256K | Nov 4, 2025 |
| moonshotai/Kimi-Linear-48B-A3B-Base | 49.1B | – | Oct 30, 2025 |
| moonshotai/Kimi-Linear-48B-A3B-Instruct | 49.1B | – | Oct 30, 2025 |
| moonshotai/Kimi-K2-Instruct-0905 | 1026B | 256K | Sep 3, 2025 |
| moonshotai/Kimi-K2-Instruct | 1026B | 128K | Jul 11, 2025 |
| moonshotai/Kimi-K2-Base | 1026B | 128K | Jul 3, 2025 |
| moonshotai/Kimi-VL-A3B-Thinking-2506 | 16.4B | 128K | Jun 21, 2025 |
| moonshotai/Kimi-Dev-72B | 72.7B | 128K | Jun 16, 2025 |
| moonshotai/Kimi-Audio-7B | 9.8B | 8K | Apr 25, 2025 |
| moonshotai/Kimi-Audio-7B-Instruct | 9.8B | 8K | Apr 25, 2025 |
| moonshotai/Kimi-VL-A3B-Thinking | 16.4B | 128K | Apr 9, 2025 |
| moonshotai/Kimi-VL-A3B-Instruct | 16.4B | 128K | Apr 9, 2025 |
Timeline
- Oct 7, 2026 · Research paper · 1 sourceFrom Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMsTo bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation.
- Oct 7, 2026 · Research paper · 1 sourceTraining Advisors for LLM Agents from Task OutcomesWe introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task.
- Oct 7, 2026 · Research paper · 1 sourceWorldBench: Evaluating LLMs on Three.js Voxel World GenerationWe present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
- Oct 7, 2026 · Research paper · 1 sourceDual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader StudyPurpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
- Oct 6, 2026 · Research paper · 1 sourceOne Step at a Time: Trading LLM Autonomy for Process PredictabilityOrganizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step.
- Oct 6, 2026 · Research paper · 1 sourceServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks.
- Oct 6, 2026 · Research paper · 1 sourceReading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language ModelsVision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
- Oct 5, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.31.0Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Oct 1, 2026 · Product / feature launch · 1 sourceunslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UXThis release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 22, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.30.0This release features 762 commits from 315 contributors (104 new)!
- Sep 18, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.20| Model | Type | PRs | Cookbook |
- Sep 9, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.17.0: Release 5.17.0Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 26, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.28.0DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Aug 8, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.17Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jul 10, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
- Jul 3, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.13.0: Release v5.13.0Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.23.0DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
- Jun 13, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.13DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
- May 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.22.0DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
- May 16, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- May 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.21.0Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).