AION
Research paperLarge Language Models · Reasoning & Planning · Robotics & Embodied AI1 source · Oct 7, 2026

Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents

Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely.

Key points

  • However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences.
  • This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model.
  • Situations may represent temporal or spatiotemporal evolution rather than only current states.
  • Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.

Sources (1)

  • [1]Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:34 AM
    Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely.
    However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences.

Extractive summary: sentences quoted from the sources.