Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.
ProofPaper ↗
Key points
- Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace.
- Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth.
- These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use.
- While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable.
Sources (1)
- [1]Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:53 PM
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.
Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
- Oct 7, 2026Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation
- Oct 7, 2026Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets
- Oct 6, 2026POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
- Sep 30, 2026huggingface/transformers v5.18.0: Release 5.18.0
- Sep 30, 2026[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU