An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under Resampling
A large body of research measures model coherence based on output variance without adequately considering competing causes.
ProofPaper ↗
Key points
- We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either.
- Even so, we find narrow finetunes score poorly.
- Inspecting inconsistencies flagged by our method, we find that model organisms from the literature display severe issues such as identity conflation, introspection failures and rationalizations.
- These findings suggest that the pathologies induced by narrow finetuning may limit what these models can tell us about coherent misaligned behaviour.
Sources (1)
- [1]An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under ResamplingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:22 PM
A large body of research measures model coherence based on output variance without adequately considering competing causes.
We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Understanding Jev, the new model everyone is talking about
- Oct 8, 2026SuperNav: An Agentic Navigation System for Any Task in Any Scene
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching