APEX: Speculate smarter, not deeper
Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth.
ProofPaper ↗
Key points
- We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation.
- APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback.
- APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens.
- We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding.
Sources (1)
- [1]APEX: Speculate smarter, not deeperarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:17 AM
Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth.
We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 22, 2026vllm-project/vllm v0.30.0
- Sep 9, 2026vllm-project/vllm v0.29.0
- Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
- Aug 26, 2026vllm-project/vllm v0.28.0
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
- Jun 29, 2026vllm-project/vllm v0.24.0