Speedbumps: Rejection Attacks on Speculative Decoding
Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass.
ProofPaper ↗
Key points
- In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle.
- We introduce two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim's prompts.
- In some cases, attacks degrade speculative decoding to the point of being slower than autoregressive decoding.
- These findings identify the draft-target interaction of speculative decoding as a realistic attack surface through which adversarial inputs can inflate inference costs.
Sources (1)
- [1]Speedbumps: Rejection Attacks on Speculative DecodingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:35 PM
Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass.
In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
- Oct 6, 2026Secure Speculative Decoding for Large Language Models
- Oct 6, 2026Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
- Oct 6, 2026DLoop: Looped Speculative Decoding
- Oct 6, 2026SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
- Oct 5, 2026vllm-project/vllm v0.31.0