Adaptive Power Sampling for LLM Reasoning
Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM).
Key points
- The goal of this work is to equip power sampling with query adaptivity.
- Theoretically, we show that the benefits of further sharpening are determined by the self-reward gap between correct and incorrect responses.
- Based on this insight, we propose Adaptive Power Sampling (APS), which adjusts the sharpening exponent on a per-query basis at test time using the relationship between answer agreement and the model's self-reward.
- Experiments across diverse reasoning tasks, including MATH500, HumanEval, and GPQA, show that APS consistently outperforms power sampling with a fixed sharpening exponent, without additional training.
Sources (1)
- [1]Adaptive Power Sampling for LLM ReasoningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:43 PM
Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM).
The goal of this work is to equip power sampling with query adaptivity.
Extractive summary: sentences quoted from the sources.