AnalysisOpinion / analysisEfficiency & Inference · Reinforcement Learning · Large Language Models1 source · Oct 11, 2026

Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.

Proof1 community thread

Key points

  • Update your llama if you haven't done so yet.
  • Main gain seems to be on prose generation.
  • Tests above ran with thinking off, ngram-mod off.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 10, 2026Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
  2. Oct 10, 2026Improve token per second without touching quant
  3. Oct 8, 2026unslothai/unsloth v0.1.905-beta: Sandboxing is here!
  4. Oct 8, 2026Microsoft Joins the Local AI Push
  5. Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
  6. Oct 7, 2026unslothai/unsloth v0.1.904-beta: Train your own Decision model

Related