What can linear attention learn from nonlinear teachers in-context?
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers.
Key points
- For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour.
- Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise.
- We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases.
- These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.
Sources (1)
- [1]What can linear attention learn from nonlinear teachers in-context?arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:24 PM
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers.
For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Leaner Transformers Can Easily Learn to Cluster
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
- Oct 6, 2026Continuous Memory Machines
- Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
- Oct 5, 2026LiquidAI/d1-omni-600M