ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment

Language models increasingly act as agents.

Key points

  • An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it.
  • On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed.
  • Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195).
  • On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  2. Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
  3. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  4. Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
  5. Oct 1, 2026nvidia/PixelUMM
  6. Aug 22, 2026sgl-project/sglang v0.5.18

Related