The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models
Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions?
ProofPaper ↗
Key points
- A retrieval-augmented model can match a document without relying on it.
- We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice.
- Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs.
- The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.
Sources (1)
- [1]The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:47 AM
Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions?
A retrieval-augmented model can match a document without relying on it.
Extractive summary: sentences quoted from the sources.