Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences.
ProofPaper ↗
Key points
- Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block.
- In multidimensional data, serialization breaks this assumption by scattering spatial neighbors across distant positions in the token sequence.
- We study how transformers overcome this routing problem in multidimensional stochastic and deterministic cellular automata, where each trajectory is generated by an unknown local rule and presented as a flattened sequence without an explicit coordinate-based spatial inductive bias.
- Attention patterns and layerwise probes align with the predicted gather-and-match computation, providing mechanistic evidence for spatial induction in trained transformers.
Sources (1)
- [1]Spatial Induction Heads: In-Context Learning of Multidimensional Cellular AutomataarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:16 PM
We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences.
Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026Continuous Memory Machines
- Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Sep 30, 2026Cloudflare/clef-flash