Mistral models
Also known as: Codestral, Devstral, Magistral, Ministral, Mistral Large, Mistral Medium, Mistral Small, Mixtral, Pixtral, Voxtral
7stories this week
8last 30 days
14all time
In the model registry
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceLocal Voice Assistant based on Qwen3.5 4B with skills/toolcalling. No dedicated GPU. Local STT+TTS
- Oct 8, 2026 · Model release · 1 sourcemistralai/Voxtral-Mini-4B-Realtime-Arabicmistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
- Oct 7, 2026 · Research paper · 1 sourceActivation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM CompressionWe propose Activation-aware Weight Tensorization (AWT), a training-free calibration wrapper that preconditions each weight matrix with a diagonal activation-derived scale before an unchanged TT/TTN solver and deploys the result with only an input-side elementwise rescaling.
- Oct 6, 2026 · Model release · 1 sourcellm-mistral 0.16Adds support for reasoning models, such as the newly released Mistral Large 4.
- Oct 6, 2026 · Product / feature launch · 1 sourceIntroducing Mistral Large 4: Le chonkIntroducing Mistral Large 4: Le chonk
- Oct 6, 2026 · Opinion / analysis · 1 sourceMistral Large 4llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
- Oct 6, 2026 · Research paper · 1 sourceOne Step at a Time: Trading LLM Autonomy for Process PredictabilityOrganizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step.
- Sep 28, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.900-beta: Laya Decision Models + LibraryRun and serve Decision Models like Laya (open-source Jev) locally
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jun 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.23.0DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
- May 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.11Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
- Apr 6, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Apr 3, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.19.0We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
- Mar 28, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10rc0Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331