NVIDIA H100
Also known as: H100, H100s
7stories this week
8last 30 days
14all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceCheaper AI tokens are driving more demand, and that's Jensen Huang's best-case scenarioData from a16z shows a Jevons paradox in the AI market: token prices keep falling, but H100 GPU rental prices hold steady or climb.
- Oct 9, 2026 · Benchmark result · 1 sourceImpactful scheduling for GPU clustersOn the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads.
- Oct 8, 2026 · Open-source release · 1 sourcehuggingface/trl v1.15.0SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
- Oct 8, 2026 · Research paper · 1 sourceLanguage Models as AI Research World ModelsAI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement.
- Oct 7, 2026 · Research paper · 2 sourcesReal Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than RecomputeWe test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
- Oct 7, 2026 · Research paper · 1 sourceYANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) MemoryTherefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning.
- Oct 7, 2026 · Model release · 2 sources[AINews] Reflection Beam - 501B-A23B American Open ModelIt’s been over a year since Reflection launched with us with big goals on coding (and hinted about their RL approach):
- Oct 2, 2026 · Opinion / analysis · 1 sourceUS arrests tech CEO accused of smuggling $300M in Nvidia chips into ChinaThe US has arrested another suspect accused of smuggling high-end computer servers containing export-controlled Nvidia chips into China.
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Aug 8, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.17Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
- Jun 10, 2026 · Model release · 1 sourceDiffusionGemma: 4x faster text generationToday, we’re introducing DiffusionGemma, an experimental open model that explores text diffusion, an exceptionally fast approach to text generation.
- May 16, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- Oct 17, 2024 · Open-source release · 1 sourcepytorch/pytorch v2.5.0: PyTorch 2.5.0 Release, SDPA CuDNN backend, Flex AttentionWe are excited to announce the release of PyTorch® 2.5!