Agentic-TTT: Training test-time policy for test-time training
To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions.
Key points
- Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems.
- By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement.
- Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions.
- On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training.
Sources (1)
- [1]Agentic-TTT: Training test-time policy for test-time trainingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:06 PM
To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions.
Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems.
Extractive summary: sentences quoted from the sources.