SanSi: A Looped Typed Decision Model for System 1.5 Thinking
We propose SanSi, which turns a pre-trained looped language model into a typed decision model.
Key points
- Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass.
- We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout.
- Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking.
- Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.
Sources (2)
- [1]SanSi: A Looped Typed Decision Model for System 1.5 ThinkingHugging Face Daily Papers · Oct 6, 12:00 AM
We propose SanSi, which turns a pre-trained looped language model into a typed decision model.
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass.
- [2]SanSi: A Looped Typed Decision Model for System 1.5 ThinkingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:28 AM · same content
Extractive summary: sentences quoted from the sources.