Nemotron
2stories this week
7last 30 days
19all time
In the model registry
| Model | Params | Context | Released |
|---|---|---|---|
| nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 | 335B | 256K | Sep 4, 2026 |
| nvidia/Nemotron-3-Labs-Ultra-Math-RL | 561B | 256K | Sep 2, 2026 |
| nvidia/Nemotron-3-Labs-Ultra-Math-SFT | 561B | 256K | Sep 2, 2026 |
| nvidia/Nemotron-3-Diarization | 99M | – | Sep 1, 2026 |
| nvidia/NVIDIA-Nemotron-Labs-Teacher-Chat | 561B | 256K | Aug 14, 2026 |
| nvidia/NVIDIA-Nemotron-Labs-Teacher-Competition-Coding | 561B | 256K | Aug 14, 2026 |
| nvidia/NVIDIA-Nemotron-Labs-Teacher-Instruction-Following | 561B | 256K | Aug 14, 2026 |
| nvidia/NVIDIA-Nemotron-Labs-Teacher-General-Reasoning | 561B | 256K | Aug 14, 2026 |
| nvidia/NVIDIA-Nemotron-Labs-Teacher-STEM | 561B | 256K | Aug 14, 2026 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 | 31.6B | 256K | Aug 5, 2026 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark | 764M | 1M | Aug 5, 2026 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash | 663M | 1M | Aug 5, 2026 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | 17.8B | 1M | Aug 4, 2026 |
| nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | 31.6B | 256K | Aug 1, 2026 |
Timeline
- Oct 7, 2026 · Opinion / analysis · 1 sourceOne Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMOStarting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.
- Oct 6, 2026 · Research paper · 1 sourceEnhancing Diffusion Language Models with Autoregressive Post-Training WeightsDiffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding.
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Oct 1, 2026 · Tutorial / explainer · 1 sourceFine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other LanguagesAutomatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data.
- Sep 30, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.18.0: Release 5.18.0Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 19, 2026 · Open-source release · 1 sourceollama/ollama v0.34.3GET /api/show now advertises each model's thinking controls and default:
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Jul 3, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.13.0: Release v5.13.0Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 13, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.13DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
- May 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.11Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
- Apr 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.20.0CUDA 13.0 default: Default CUDA wheel on PyPI and vllm/vllm-openai:v0.20.0 image switched to CUDA 13.0; architecture lists and build-args cleaned up (#39878), and CUDA bumped to 13.0.2 to match PyTorch 2.11.0 (#40669).
- Apr 6, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Apr 3, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.19.0We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
- Mar 28, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10rc0Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Jan 9, 2026 · Open-source release · 1 sourcesgl-project/sglang gateway-v0.3.1: Release Gateway-v0.3.1We're excited to announce SMG v0.3.1 – a game-changing release with 10-12x performance improvement and 99% memory reduction in cache-aware routing, plus enterprise-grade security!
- Jan 1, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.7[SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered