Hardware

NVIDIA H100

Also known as: H100, H100s

7stories this week
8last 30 days
14all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Cheaper AI tokens are driving more demand, and that's Jensen Huang's best-case scenario
    Data from a16z shows a Jevons paradox in the AI market: token prices keep falling, but H100 GPU rental prices hold steady or climb.
  2. Oct 9, 2026 · Benchmark result · 1 source
    Impactful scheduling for GPU clusters
    On the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads.
  3. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  4. Oct 8, 2026 · Research paper · 1 source
    Language Models as AI Research World Models
    AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement.
  5. Oct 7, 2026 · Research paper · 2 sources
    Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
    We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
  6. Oct 7, 2026 · Research paper · 1 source
    YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory
    Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning.
  7. Oct 7, 2026 · Model release · 2 sources
    [AINews] Reflection Beam - 501B-A23B American Open Model
    It’s been over a year since Reflection launched with us with big goals on coding (and hinted about their RL approach):
  8. Oct 2, 2026 · Opinion / analysis · 1 source
    US arrests tech CEO accused of smuggling $300M in Nvidia chips into China
    The US has arrested another suspect accused of smuggling high-end computer servers containing export-controlled Nvidia chips into China.
  9. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  10. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  11. Aug 8, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.17
    Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
  12. Jun 10, 2026 · Model release · 1 source
    DiffusionGemma: 4x faster text generation
    Today, we’re introducing DiffusionGemma, an experimental open model that explores text diffusion, an exceptionally fast approach to text generation.
  13. May 16, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12
    DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
  14. Oct 17, 2024 · Open-source release · 1 source
    pytorch/pytorch v2.5.0: PyTorch 2.5.0 Release, SDPA CuDNN backend, Flex Attention
    We are excited to announce the release of PyTorch® 2.5!

Often appears with