AION
Research paperReinforcement Learning1 source · Oct 6, 2026

VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents

We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic.

Key points

  • Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token.
  • This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response?
  • VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates.
  • These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient.

Sources (1)

  • [1]VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:15 PM
    We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic.
    Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token.

Extractive summary: sentences quoted from the sources.