Reminder: try probabilistic MTP if you missed it. Decode +14% on prose
Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.
Proof1 community thread
Key points
- Update your llama if you haven't done so yet.
- Main gain seems to be on prose generation.
- Tests above ran with thinking off, ngram-mod off.
Sources (1)
- [1]Reminder: try probabilistic MTP if you missed it. Decode +14% on proser/LocalLLaMA (top, daily) · Oct 11, 09:20 AM
Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.
Update your llama if you haven't done so yet.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 10, 2026Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
- Oct 10, 2026Improve token per second without touching quant
- Oct 8, 2026unslothai/unsloth v0.1.905-beta: Sandboxing is here!
- Oct 8, 2026Microsoft Joins the Local AI Push
- Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
- Oct 7, 2026unslothai/unsloth v0.1.904-beta: Train your own Decision model