AION
Hardware

NVIDIA Blackwell

Also known as: B200, B300, Blackwell, GB200, GB300

7stories this week
12last 30 days
34all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens
    Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€
  2. Oct 9, 2026 · Opinion / analysis · 1 source
    Impactful scheduling for GPU clusters
    On the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads.
  3. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  4. Oct 8, 2026 · Research paper · 1 source
    PageWeaver: KV-Guided Query Unions for Sparse Attention
    Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work.
  5. Oct 7, 2026 · Research paper · 1 source
    The Missing Fourth Term for the Emulation Tensor Memory Equilibrium (TME) Model: The Residue Deconstruction Cost
    The Tensor-Memory Equilibrium (TME) model of "FP8 is All You Need (Part 1)" calculates the execution time of Ozaki Scheme II emulation of fp64 as the maximum of a tensor-core term and a High-Bandwidth Memory (HBM) traffic term, plus a per-output reconstruction term.
  6. Oct 7, 2026 · Opinion / analysis · 1 source
    The Machines that Make the Machines
    However, the process of assembling GB300 trays requires skilled physical labor in factories across the world.
  7. Oct 6, 2026 · Product / feature launch · 1 source
    Introducing Mistral Large 4: Le chonk
    Introducing Mistral Large 4: Le chonk
  8. Oct 1, 2026 · Product / feature launch · 1 source
    unslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UX
    This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
  9. Sep 29, 2026 · Open-source release · 1 source
    NVIDIA/TensorRT-LLM v1.3.0rc29
    Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
  10. Sep 28, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
    Run and serve Decision Models like Laya (open-source Jev) locally
  11. Sep 22, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.30.0
    This release features 762 commits from 315 contributors (104 new)!
  12. Sep 18, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.20
    | Model | Type | PRs | Cookbook |
  13. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  14. Aug 26, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.28.0
    DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
  15. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  16. Aug 8, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.17
    Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
  17. Jul 25, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.16
    DSpark: confidence-driven speculative decoding: A new speculative algorithm.
  18. Jul 10, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15
    GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
  19. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  20. Jun 18, 2026 · Open-source release · 1 source
    pytorch/pytorch v2.12.1: PyTorch 2.12.1 Release, bug fix release
    This release is meant to fix the following regressions and silent correctness issues:
  21. Jun 13, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.13
    DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
  22. May 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.22.0
    DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
  23. May 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12.post1
    v0.5.12.post1 is a stability patch on top of v0.5.12.
  24. May 16, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12
    DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
  25. May 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.21.0
    Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
  26. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  27. Apr 3, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.19.0
    We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
  28. Mar 31, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.18.1
    This is a patch release on top of v0.18.0 to address a few issues:
  29. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  30. Mar 23, 2026 · Open-source release · 1 source
    pytorch/pytorch v2.11.0: PyTorch 2.11.0 Release
    <strong>FlexAttention</strong> now has a <strong>FlashAttention-4</strong> backend on <strong>Hopper</strong> and <strong>Blackwell</strong> GPUs

Often appears with