BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models
We introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition.
Key points
- A language model may revise the same proposition after receiving genuinely relevant evidence or after receiving directional user pressure that adds no relevant fact.
- BeliefScope crosses Evidence and Pressure with factor-specific local controls and measures response changes through probability reports, categorical judgments, and action recommendations on channel-appropriate scales.
- To determine when these observable contrasts support reliable attribution, we evaluate the observation design under controlled synthetic conditions.
- Across a 36-family Qwen/Llama study, with targeted 12-family checks that also include Gemma3-12B, the resulting profiles show substantial evaluation-context dependence: broad model-level differences can change under matched controls, decoding, or response interfaces, while some narrower within-model patterns remain stable.
Sources (1)
- [1]BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:11 AM
We introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition.
A language model may revise the same proposition after receiving genuinely relevant evidence or after receiving directional user pressure that adds no relevant fact.
Extractive summary: sentences quoted from the sources.