Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models
Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available?
ProofPaper ↗
Key points
- We show the answer is encoded in the spectral structure of pretrained weights.
- We prove that the OOD accuracy gap is bounded by how tightly the source representations concentrate.
- The selection does not leak the target: for each model family outside the matrix we logged the cell, metric and sign before running its OOD evaluation, and the predicted direction held in every case: EEG, genomic and protein.
- The diagnostic operates on released weights alone, so OOD robustness becomes checkable at model-selection time, before data or compute is committed to a target domain.
Sources (1)
- [1]Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:04 AM
Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available?
We show the answer is encoded in the spectral structure of pretrained weights.
Extractive summary: sentences quoted from the sources.