AION
Research paperReasoning & Planning1 source · Oct 8, 2026

ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry

We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.

Key points

  • Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student.
  • Reflection offers a way to derive explicit error diagnoses and revision guidance from self-generated attempts, yet existing reflection-based methods often combine it with reference information, rich task feedback, or persistent memory.
  • An outcome-aware selection and weighting strategy distinguishes initially correct, reflection-corrected, and unresolved examples, assigning separate weights to their category-normalized distillation losses.
  • Through on-policy distillation, the student matches the teacher's context-conditioned token-level predictive distributions at prefixes of its own rollouts, transferring the benefits of iterative correction while retaining single-pass inference.

Sources (1)

  • [1]ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:58 AM
    We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
    Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student.

Extractive summary: sentences quoted from the sources.