The Standardization Trap: Certifying Joint Label Processing in Tabular Foundation Models
Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label.
ProofPaper ↗
Key points
- Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map.
- We propose two certificates that depend only on predictions at standardized labels and can reject two distinct explanations: fixed-weight prediction and sums of independent nonlinear label transformations.
- Across the five public TFMs that we evaluate, our certificates show that changing one context label alters how other labels influence the prediction, a behavior we call joint processing.
- We further find that joint processing emerges with training and that attention scores carry most of the measured interaction.
Sources (1)
- [1]The Standardization Trap: Certifying Joint Label Processing in Tabular Foundation ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 01:18 PM
Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label.
Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026TICDA: Tabular In-Context Data Attribution
- Oct 6, 2026Continuous Memory Machines
- Oct 6, 2026Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
- Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
- Sep 29, 2026In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks