Repository / library

vLLM

15stories this week
21last 30 days
44all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Is a second MSc worth it? [D]
    I’m an AI research engineer based in Africa with about 3 YoEs in RL, LLMs, post-training, and systems efficiency (CUDA/vLLM).
  2. Oct 11, 2026 · Benchmark result · 1 source
    Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
    Speculative decoding drops a demo model from roughly 58 tokens per second to 16 instead of speeding it up.
  3. Oct 10, 2026 · Opinion / analysis · 1 source
    48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)
    Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
  4. Oct 10, 2026 · Model release · 4 sources
    Microsoft's Decision-1 model enters the fast-growing AI decision model race
    With Decision-1, Microsoft enters the growing decision model space.
  5. Oct 10, 2026 · Opinion / analysis · 1 source
    Strata with Qwen3.8 Flash Next UD-Q4_K_XL
    Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
  6. Oct 10, 2026 · Opinion / analysis · 1 source
    [AINews] TypeSafe/Jev at >$100M ARR, $7.5B valuation 3 weeks after launch
    As you can see in the AINews X recap section below, everyone on earth has cloned the Jev API, but only one company can ever create the category.
  7. Oct 9, 2026 · Model release · 2 sources
    ConwayResearch/Underdog-Saluki-27B-1.0
    ConwayResearch published the model Underdog-Saluki-27B-1.0 on Hugging Face.
  8. Oct 8, 2026 · Model release · 1 source
    mistralai/Voxtral-Mini-4B-Realtime-Arabic
    mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
  9. Oct 7, 2026 · Research paper · 2 sources
    Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
    We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
  10. Oct 7, 2026 · Research paper · 1 source
    Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing
    Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state.
  11. Oct 6, 2026 · Open-source release · 1 source
    huggingface/trl v1.14.2
    Patch release fixing two cases of silently wrong training and three crashes.
  12. Oct 6, 2026 · Research paper · 1 source
    SPIN: Shadow Predictive Indexer for Sparse Attention
    We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead.
  13. Oct 6, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.903-beta: New Browser + Voice Cloning
    This release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
  14. Oct 6, 2026 · Research paper · 1 source
    APEX: Speculate smarter, not deeper
    Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth.
  15. Oct 5, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.31.0
    Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
  16. Oct 2, 2026 · Open-source release · 1 source
    ray-project/ray ray-2.59.0: Ray-2.59.0
    Ray Data LLM & Ray Serve LLM are GA/Stable: the LLM APIs graduate to general availability this release (\#65194), alongside an upgrade to vLLM 0.27.0 (\#65351).
  17. Oct 2, 2026 · Model release · 1 source
    Aleph-Alpha/Kolibri-1
    Aleph-Alpha published the model Kolibri-1 on Hugging Face.
  18. Sep 30, 2026 · Open-source release · 1 source
    huggingface/transformers v5.18.0: Release 5.18.0
    Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
  19. Sep 30, 2026 · Research paper · 1 source
    SparseEngine: Sparse-First Inference Engine
    We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure.
  20. Sep 29, 2026 · Open-source release · 1 source
    huggingface/trl v1.14.1
  21. Sep 22, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.30.0
    This release features 762 commits from 315 contributors (104 new)!
  22. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  23. Aug 26, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.28.0
    DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
  24. Aug 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.1
    Support quantized DSpark Markov heads (#50424)
  25. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  26. Jul 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.26.0
    New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
  27. Jul 15, 2026 · Open-source release · 1 source
    huggingface/transformers v5.14.0: Release v5.14.0
    Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
  28. Jul 14, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.1
    Avoid blocking model launching when no system FFmpeg is available for TorchCodec (#47888).
  29. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  30. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).

Often appears with