AMD
3stories this week
6last 30 days
22all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceAMD Reportedly Raises GDDR6 Prices for Board PartnersIs an AMD price incoming as well?
- Oct 11, 2026 · Tutorial / explainer · 1 sourceBuilding a 4x R9700 setup for a 10 person startupJust wanted to share a build I am doing for a client.
- Oct 10, 2026 · Opinion / analysis · 1 sourceImprove token per second without touching quantSpent the past month tweaking and experimenting with many different numbers to achieve 30tps.
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Oct 1, 2026 · Product / feature launch · 1 sourceunslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UXThis release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
- Sep 30, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.18.0: Release 5.18.0Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 26, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.28.0DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Aug 8, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.17Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jul 25, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.16DSpark: confidence-driven speculative decoding: A new speculative algorithm.
- Jul 10, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 13, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.13DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
- Jun 5, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.22.1v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions.
- May 16, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- May 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.21.0Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
- May 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.11Speculative Decoding V2 by default: Spec V2 (with overlap scheduling to hide CPU overhead) is now the default, materially reducing per-step CPU cost for EAGLE/MTP/DFLASH paths: #21062
- Jan 23, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.8Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
- Jan 1, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.7[SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered