Multi-token prediction
Also known as: MTP
9stories this week
13last 30 days
40all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceReminder: try probabilistic MTP if you missed it. Decode +14% on proseOptimal draft-n-max / draft-p-min seem to be in line with greedy sampling.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceRunning Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boardsThis will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceOMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
- Oct 10, 2026 · Opinion / analysis · 1 source48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
- Oct 10, 2026 · Opinion / analysis · 1 sourceQwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernelI've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
- Oct 10, 2026 · Opinion / analysis · 1 sourceImprove token per second without touching quantSpent the past month tweaking and experimenting with many different numbers to achieve 30tps.
- Oct 10, 2026 · Opinion / analysis · 1 sourceStrata with Qwen3.8 Flash Next UD-Q4_K_XLMost of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
- Oct 6, 2026 · Research paper · 1 sourceDLoop: Looped Speculative DecodingSpeculative decoding accelerates autoregressive generation in large language models.
- Oct 5, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.31.0Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
- Oct 4, 2026 · Model release · 1 sourcenerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2nerkyor published the model Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 on Hugging Face.
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 22, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.30.0This release features 762 commits from 315 contributors (104 new)!
- Sep 9, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.17.0: Release 5.17.0Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 26, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.16.0: Release: v5.16.0Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
- Aug 26, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.28.0DeepSeek V4: sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding (#51538), joined by AMD Quark NVFP4 support (#47972), reasoning-effort prompts and mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212).
- Aug 19, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.15.1: Patch release: v5.15.1This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter.
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jul 25, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.16DSpark: confidence-driven speculative decoding: A new speculative algorithm.
- Jul 15, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.14.0: Release v5.14.0Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jul 10, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.23.0DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
- Jun 13, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.13DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
- Jun 9, 2026 · Model release · 1 sourceIntroducing Gemma 4 12B: a unified, encoder-free multimodal modelIntroducing Gemma 4 12B: a unified, encoder-free multimodal model