ResearchResearch paperEfficiency & Inference1 source · Oct 6, 2026

APEX: Speculate smarter, not deeper

Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth.

Key points

  • We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation.
  • APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback.
  • APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens.
  • We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding.

Sources (1)

  • [1]APEX: Speculate smarter, not deeper
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:17 AM
    Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth.
    We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 22, 2026vllm-project/vllm v0.30.0
  2. Sep 9, 2026vllm-project/vllm v0.29.0
  3. Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
  4. Aug 26, 2026vllm-project/vllm v0.28.0
  5. Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
  6. Jun 29, 2026vllm-project/vllm v0.24.0

Related