AION
Model

Nemotron

2stories this week
7last 30 days
19all time

In the model registry

Timeline

  1. Oct 7, 2026 · Opinion / analysis · 1 source
    One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
    Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.
  2. Oct 6, 2026 · Research paper · 1 source
    Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
    Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding.
  3. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  4. Oct 1, 2026 · Tutorial / explainer · 1 source
    Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages
    Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data.
  5. Sep 30, 2026 · Open-source release · 1 source
    huggingface/transformers v5.18.0: Release 5.18.0
    Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
  6. Sep 29, 2026 · Open-source release · 1 source
    NVIDIA/TensorRT-LLM v1.3.0rc29
    Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
  7. Sep 19, 2026 · Open-source release · 1 source
    ollama/ollama v0.34.3
    GET /api/show now advertises each model's thinking controls and default:
  8. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  9. Jul 3, 2026 · Open-source release · 1 source
    huggingface/transformers v5.13.0: Release v5.13.0
    Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
  10. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  11. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  12. Jun 13, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.13
    DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
  13. May 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.11
    Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
  14. Apr 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.20.0
    CUDA 13.0 default: Default CUDA wheel on PyPI and vllm/vllm-openai:v0.20.0 image switched to CUDA 13.0; architecture lists and build-args cleaned up (#39878), and CUDA bumped to 13.0.2 to match PyTorch 2.11.0 (#40669).
  15. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  16. Apr 3, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.19.0
    We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
  17. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  18. Jan 9, 2026 · Open-source release · 1 source
    sgl-project/sglang gateway-v0.3.1: Release Gateway-v0.3.1
    We're excited to announce SMG v0.3.1 – a game-changing release with 10-12x performance improvement and 99% memory reduction in cache-aware routing, plus enterprise-grade security!
  19. Jan 1, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.7
    [SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered

Often appears with