VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic.
Key points
- Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token.
- This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response?
- VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates.
- These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient.
Sources (1)
- [1]VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:15 PM
We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic.
Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token.
Extractive summary: sentences quoted from the sources.