ResearchResearch paperInterpretability1 source · Oct 6, 2026

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict.

Key points

  • Probes are the workhorse of interpretability.
  • If a model's hidden states predict a variable, the model is said to represent it.
  • We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it.
  • The single-cell foundation model scGPT encodes biological variability only partially.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 5, 2026LiquidAI/d1-omni-600M
  2. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  3. Sep 30, 2026Cloudflare/clef-flash
  4. Sep 29, 2026microsoft/AesCode-32B
  5. Sep 29, 2026Language Models for Text Classification: From Bag-of-Words to Jev
  6. Aug 26, 2026vllm-project/vllm v0.28.0

Related