AION
Research paperLarge Language Models · Interpretability1 source · Oct 7, 2026

What can linear attention learn from nonlinear teachers in-context?

Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers.

Key points

  • For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour.
  • Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise.
  • We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases.
  • These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.

Sources (1)

  • [1]What can linear attention learn from nonlinear teachers in-context?
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:24 PM
    Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers.
    For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Leaner Transformers Can Easily Learn to Cluster
  2. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  3. Oct 6, 2026Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
  4. Oct 6, 2026Continuous Memory Machines
  5. Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
  6. Oct 5, 2026LiquidAI/d1-omni-600M

Related