ResearchResearch paperLarge Language Models · Interpretability1 source · Oct 6, 2026

The Standardization Trap: Certifying Joint Label Processing in Tabular Foundation Models

Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label.

Key points

  • Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map.
  • We propose two certificates that depend only on predictions at standardized labels and can reject two distinct explanations: fixed-weight prediction and sums of independent nonlinear label transformations.
  • Across the five public TFMs that we evaluate, our certificates show that changing one context label alters how other labels influence the prediction, a behavior we call joint processing.
  • We further find that joint processing emerges with training and that attention scores carry most of the measured interaction.

Sources (1)

  • [1]The Standardization Trap: Certifying Joint Label Processing in Tabular Foundation Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 01:18 PM
    Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label.
    Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026TICDA: Tabular In-Context Data Attribution
  2. Oct 6, 2026Continuous Memory Machines
  3. Oct 6, 2026Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
  4. Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
  5. Sep 29, 2026In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

Related