NVIDIA Blackwell
Also known as: B200, B300, Blackwell, GB200, GB300
7stories this week
12last 30 days
34all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceI trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokensHiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€
- Oct 9, 2026 · Opinion / analysis · 1 sourceImpactful scheduling for GPU clustersOn the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads.
- Oct 8, 2026 · Open-source release · 1 sourcehuggingface/trl v1.15.0SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
- Oct 8, 2026 · Research paper · 1 sourcePageWeaver: KV-Guided Query Unions for Sparse AttentionDynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work.
- Oct 7, 2026 · Research paper · 1 sourceThe Missing Fourth Term for the Emulation Tensor Memory Equilibrium (TME) Model: The Residue Deconstruction CostThe Tensor-Memory Equilibrium (TME) model of "FP8 is All You Need (Part 1)" calculates the execution time of Ozaki Scheme II emulation of fp64 as the maximum of a tensor-core term and a High-Bandwidth Memory (HBM) traffic term, plus a per-output reconstruction term.
- Oct 7, 2026 · Opinion / analysis · 1 sourceThe Machines that Make the MachinesHowever, the process of assembling GB300 trays requires skilled physical labor in factories across the world.
- Oct 6, 2026 · Product / feature launch · 1 sourceIntroducing Mistral Large 4: Le chonkIntroducing Mistral Large 4: Le chonk
- Oct 1, 2026 · Product / feature launch · 1 sourceunslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UXThis release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 28, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.900-beta: Laya Decision Models + LibraryRun and serve Decision Models like Laya (open-source Jev) locally
- Sep 22, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.30.0This release features 762 commits from 315 contributors (104 new)!
- Sep 18, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.20| Model | Type | PRs | Cookbook |
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 26, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.28.0DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Aug 8, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.17Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
- Jul 25, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.16DSpark: confidence-driven speculative decoding: A new speculative algorithm.
- Jul 10, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 18, 2026 · Open-source release · 1 sourcepytorch/pytorch v2.12.1: PyTorch 2.12.1 Release, bug fix releaseThis release is meant to fix the following regressions and silent correctness issues:
- Jun 13, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.13DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
- May 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.22.0DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
- May 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12.post1v0.5.12.post1 is a stability patch on top of v0.5.12.
- May 16, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- May 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.21.0Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
- Apr 6, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Apr 3, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.19.0We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
- Mar 31, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.18.1This is a patch release on top of v0.18.0 to address a few issues:
- Mar 28, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10rc0Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Mar 23, 2026 · Open-source release · 1 sourcepytorch/pytorch v2.11.0: PyTorch 2.11.0 Release<strong>FlexAttention</strong> now has a <strong>FlashAttention-4</strong> backend on <strong>Hopper</strong> and <strong>Blackwell</strong> GPUs