AION
Model

Mistral models

Also known as: Codestral, Devstral, Magistral, Ministral, Mistral Large, Mistral Medium, Mistral Small, Mixtral, Pixtral, Voxtral

7stories this week
8last 30 days
14all time

In the model registry

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Local Voice Assistant based on Qwen3.5 4B with skills/toolcalling. No dedicated GPU. Local STT+TTS
  2. Oct 8, 2026 · Model release · 1 source
    mistralai/Voxtral-Mini-4B-Realtime-Arabic
    mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
  3. Oct 7, 2026 · Research paper · 1 source
    Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression
    We propose Activation-aware Weight Tensorization (AWT), a training-free calibration wrapper that preconditions each weight matrix with a diagonal activation-derived scale before an unchanged TT/TTN solver and deploys the result with only an input-side elementwise rescaling.
  4. Oct 6, 2026 · Model release · 1 source
    llm-mistral 0.16
    Adds support for reasoning models, such as the newly released Mistral Large 4.
  5. Oct 6, 2026 · Product / feature launch · 1 source
    Introducing Mistral Large 4: Le chonk
    Introducing Mistral Large 4: Le chonk
  6. Oct 6, 2026 · Opinion / analysis · 1 source
    Mistral Large 4
    llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
  7. Oct 6, 2026 · Research paper · 1 source
    One Step at a Time: Trading LLM Autonomy for Process Predictability
    Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step.
  8. Sep 28, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
    Run and serve Decision Models like Laya (open-source Jev) locally
  9. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  10. Jun 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.23.0
    DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
  11. May 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.11
    Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
  12. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  13. Apr 3, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.19.0
    We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
  14. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331

Often appears with