How to train your model organism
We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities.
ProofPaper ↗
Key points
- Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques.
- We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations).
- We introduce a multi-objective training approach based on model merging to train more realistic model organisms.
- Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness.
Sources (1)
- [1]How to train your model organismarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:02 PM
We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities.
Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching
- Oct 6, 2026AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 30, 2026[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU
- Sep 29, 2026The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models