ResearchResearch paperRobotics & Embodied AI · Efficiency & Inference · Large Language Models1 source · Oct 6, 2026

Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata

We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences.

Key points

  • Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block.
  • In multidimensional data, serialization breaks this assumption by scattering spatial neighbors across distant positions in the token sequence.
  • We study how transformers overcome this routing problem in multidimensional stochastic and deterministic cellular automata, where each trajectory is generated by an unknown local rule and presented as a flattened sequence without an explicit coordinate-based spatial inductive bias.
  • Attention patterns and layerwise probes align with the predicted gather-and-match computation, providing mechanistic evidence for spatial induction in trained transformers.

Sources (1)

  • [1]Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:16 PM
    We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences.
    Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  2. Oct 6, 2026Continuous Memory Machines
  3. Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
  4. Oct 5, 2026LiquidAI/d1-omni-600M
  5. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  6. Sep 30, 2026Cloudflare/clef-flash

Related