How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict.
ProofPaper ↗
Key points
- Probes are the workhorse of interpretability.
- If a model's hidden states predict a variable, the model is said to represent it.
- We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it.
- The single-cell foundation model scGPT encodes biological variability only partially.
Sources (1)
- [1]How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability ProbingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:32 PM
We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict.
Probes are the workhorse of interpretability.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Sep 30, 2026Cloudflare/clef-flash
- Sep 29, 2026microsoft/AesCode-32B
- Sep 29, 2026Language Models for Text Classification: From Bag-of-Words to Jev
- Aug 26, 2026vllm-project/vllm v0.28.0