Can Jev be Your Q or Policy in Reinforcement Learning?
Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly.
Key points
- Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass.
- Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains.
- We then construct algorithms that employ Jev at three of these positions, as a reference policy, an exploration judge, and a replay rater, to improve sample efficiency, exploration, and learning performance.
- To our knowledge, we present the first use of Jev within the RL learning process and establish a frozen decision model as a usable component of RL training, inviting further exploration of how Jev and other advanced decision models can improve RL.
Sources (1)
- [1]Can Jev be Your Q or Policy in Reinforcement Learning?arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:02 AM
Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly.
Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass.
Extractive summary: sentences quoted from the sources.