AION
Technique

Multi-token prediction

Also known as: MTP

9stories this week
13last 30 days
40all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Reminder: try probabilistic MTP if you missed it. Decode +14% on prose
    Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.
  2. Oct 11, 2026 · Tutorial / explainer · 1 source
    Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards
    This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
  3. Oct 11, 2026 · Tutorial / explainer · 1 source
    OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!
    I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
  4. Oct 10, 2026 · Opinion / analysis · 1 source
    48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)
    Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
  5. Oct 10, 2026 · Opinion / analysis · 1 source
    Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
    I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
  6. Oct 10, 2026 · Opinion / analysis · 1 source
    Improve token per second without touching quant
    Spent the past month tweaking and experimenting with many different numbers to achieve 30tps.
  7. Oct 10, 2026 · Opinion / analysis · 1 source
    Strata with Qwen3.8 Flash Next UD-Q4_K_XL
    Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
  8. Oct 6, 2026 · Research paper · 1 source
    DLoop: Looped Speculative Decoding
    Speculative decoding accelerates autoregressive generation in large language models.
  9. Oct 5, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.31.0
    Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
  10. Oct 4, 2026 · Model release · 1 source
    nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
    nerkyor published the model Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 on Hugging Face.
  11. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  12. Sep 29, 2026 · Open-source release · 1 source
    NVIDIA/TensorRT-LLM v1.3.0rc29
    Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
  13. Sep 22, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.30.0
    This release features 762 commits from 315 contributors (104 new)!
  14. Sep 9, 2026 · Open-source release · 1 source
    huggingface/transformers v5.17.0: Release 5.17.0
    Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
  15. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  16. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  17. Aug 26, 2026 · Open-source release · 1 source
    huggingface/transformers v5.16.0: Release: v5.16.0
    Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
  18. Aug 26, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.28.0
    DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
  19. Aug 19, 2026 · Open-source release · 1 source
    huggingface/transformers v5.15.1: Patch release: v5.15.1
    This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter.
  20. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  21. Jul 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.26.0
    New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
  22. Jul 25, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.16
    DSpark: confidence-driven speculative decoding: A new speculative algorithm.
  23. Jul 15, 2026 · Open-source release · 1 source
    huggingface/transformers v5.14.0: Release v5.14.0
    Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
  24. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  25. Jul 10, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15
    GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
  26. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  27. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  28. Jun 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.23.0
    DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
  29. Jun 13, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.13
    DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
  30. Jun 9, 2026 · Model release · 1 source
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Often appears with