ResearchResearch paperLarge Language Models · Reinforcement Learning1 source · Oct 6, 2026

Constraint Tree Exploration for Learning from Language Feedback

Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements.

Key points

  • We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions.
  • We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it.
  • With reliable identification, TRACE-Identification can replace this dependence by $K/p{ext}$, where $K$ is the number of latent constraints and $p{ext}$ lower-bounds the probability of extracting a missing true constraint from informative feedback.
  • We evaluate TRACE across six language-feedback tasks.

Sources (1)

  • [1]Constraint Tree Exploration for Learning from Language Feedback
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:57 PM
    Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements.
    We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions.

Extractive summary: sentences quoted from the sources.

Related