AION
Research paperReinforcement Learning · Training & Scaling1 source · Oct 8, 2026

Agentic-TTT: Training test-time policy for test-time training

To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions.

Key points

  • Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems.
  • By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement.
  • Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions.
  • On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training.

Sources (1)

  • [1]Agentic-TTT: Training test-time policy for test-time training
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:06 PM
    To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions.
    Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems.

Extractive summary: sentences quoted from the sources.