Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Language models increasingly act as agents.
ProofPaper ↗
Key points
- An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it.
- On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed.
- Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195).
- On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that.
Sources (1)
- [1]Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral JudgmentarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:52 PM
Language models increasingly act as agents.
An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
- Oct 1, 2026nvidia/PixelUMM
- Aug 22, 2026sgl-project/sglang v0.5.18