Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps.
ProofPaper ↗
Key points
- To address these limitations, we introduce LoopOPD, a cross-loop on-policy distillation framework that uses additional recurrent computation within a LoopLM as its own source of supervision.
- We further propose Dynamic LoopOPD (D-LoopOPD), which continually refreshes the terminal loop teacher as the shared model parameters are updated, enabling recurrent self-improvement.
- Experiments on Ouro-Thinking models show that LoopOPD improves mathematical reasoning, while D-LoopOPD yields further gains through dynamic teacher updates.
- Despite being trained only on mathematical data, the resulting models also improve on general reasoning and code generation benchmarks, demonstrating that recurrent computation can serve as an effective source of supervision for LoopLMs. Our code and model checkpoints will be released upon acceptance.
Sources (1)
- [1]Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:54 AM
Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps.
To address these limitations, we introduce LoopOPD, a cross-loop on-policy distillation framework that uses additional recurrent computation within a LoopLM as its own source of supervision.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
- Oct 7, 2026Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
- Oct 7, 2026On-Policy Distillation Teaches New Skills but Not New Knowledge
- Oct 7, 2026UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more