ResearchResearch paperEfficiency & Inference1 source · Oct 7, 2026

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel.

Key points

  • Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted.
  • In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process.
  • This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds.
  • We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts.

Sources (1)

  • [1]Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:56 PM
    Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel.
    Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 7, 2026Q-Learning with Scalar Adjoint Matching
  3. Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
  4. Oct 6, 2026SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
  5. Oct 5, 2026vllm-project/vllm v0.31.0
  6. Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK

Related