vLLM
15stories this week
21last 30 days
44all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceIs a second MSc worth it? [D]I’m an AI research engineer based in Africa with about 3 YoEs in RL, LLMs, post-training, and systems efficiency (CUDA/vLLM).
- Oct 11, 2026 · Benchmark result · 1 sourceAgents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai TechnologiesSpeculative decoding drops a demo model from roughly 58 tokens per second to 16 instead of speeding it up.
- Oct 10, 2026 · Opinion / analysis · 1 source48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
- Oct 10, 2026 · Model release · 4 sourcesMicrosoft's Decision-1 model enters the fast-growing AI decision model raceWith Decision-1, Microsoft enters the growing decision model space.
- Oct 10, 2026 · Opinion / analysis · 1 sourceStrata with Qwen3.8 Flash Next UD-Q4_K_XLMost of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
- Oct 10, 2026 · Opinion / analysis · 1 source[AINews] TypeSafe/Jev at >$100M ARR, $7.5B valuation 3 weeks after launchAs you can see in the AINews X recap section below, everyone on earth has cloned the Jev API, but only one company can ever create the category.
- Oct 9, 2026 · Model release · 2 sourcesConwayResearch/Underdog-Saluki-27B-1.0ConwayResearch published the model Underdog-Saluki-27B-1.0 on Hugging Face.
- Oct 8, 2026 · Model release · 1 sourcemistralai/Voxtral-Mini-4B-Realtime-Arabicmistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
- Oct 7, 2026 · Research paper · 2 sourcesReal Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than RecomputeWe test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
- Oct 7, 2026 · Research paper · 1 sourceReproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression TestingReproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state.
- Oct 6, 2026 · Open-source release · 1 sourcehuggingface/trl v1.14.2Patch release fixing two cases of silently wrong training and three crashes.
- Oct 6, 2026 · Research paper · 1 sourceSPIN: Shadow Predictive Indexer for Sparse AttentionWe propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead.
- Oct 6, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.903-beta: New Browser + Voice CloningThis release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
- Oct 6, 2026 · Research paper · 1 sourceAPEX: Speculate smarter, not deeperSpeculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth.
- Oct 5, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.31.0Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
- Oct 2, 2026 · Open-source release · 1 sourceray-project/ray ray-2.59.0: Ray-2.59.0Ray Data LLM & Ray Serve LLM are GA/Stable: the LLM APIs graduate to general availability this release (\#65194), alongside an upgrade to vLLM 0.27.0 (\#65351).
- Oct 2, 2026 · Model release · 1 sourceAleph-Alpha/Kolibri-1Aleph-Alpha published the model Kolibri-1 on Hugging Face.
- Sep 30, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.18.0: Release 5.18.0Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
- Sep 30, 2026 · Research paper · 1 sourceSparseEngine: Sparse-First Inference EngineWe present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.
- Sep 29, 2026 · Open-source release · 1 sourcehuggingface/trl v1.14.1
- Sep 22, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.30.0This release features 762 commits from 315 contributors (104 new)!
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Aug 26, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.28.0DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
- Aug 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.1Support quantized DSpark Markov heads (#50424)
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jul 15, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.14.0: Release v5.14.0Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
- Jul 14, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.1Avoid blocking model launching when no system FFmpeg is available for TorchCodec (#47888).
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).