Chain-of-thought
Also known as: CoT, chain of thought
19stories this week
24last 30 days
26all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceEvery Model That Can Be Run On 10-16GB VRAM RankedIt's been 4 years since c.ai first hallucination model, and yet, we're nowhere good enough at LLMs in terms of spontaneity/interesting hallucination features.
- Oct 8, 2026 · Research paper · 1 sourceCited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought FaithfulnessLarge language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.
- Oct 8, 2026 · Research paper · 1 sourceFrom Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped TransformersWe introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.
- Oct 8, 2026 · Research paper · 1 sourceBeyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-SpeechNatural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
- Oct 8, 2026 · Research paper · 2 sourcesSpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent TracesSpatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
- Oct 8, 2026 · Research paper · 1 sourceDeception by Omission: Language Models Knowingly Hide Their MistakesLarge language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed.
- Oct 8, 2026 · Research paper · 1 sourceLapras: Latent Reasoning for Time Series Language ModelsTime Series Language Models (TSLMs) offer a promising path toward time series understanding by reasoning over temporal signals and producing natural language answers and explanations.
- Oct 7, 2026 · Research paper · 1 sourceSPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level AdvantagesVision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
- Oct 7, 2026 · Research paper · 1 sourceKDFP: A first-principles approach to knowledge distillation in large language modelsKnowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models.
- Oct 7, 2026 · Research paper · 1 sourceEngramEdit: Decoupled Knowledge Updates in LLMs through Conditional MemoryWe propose EngramEdit for decoupled knowledge updates through conditional memory.
- Oct 7, 2026 · Research paper · 1 sourceReasoning-Token Spikes Under Prompted Untruthful Responding in Large Language ModelsMonitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.
- Oct 7, 2026 · Research paper · 1 sourceExplicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous DrivingVision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
- Oct 7, 2026 · Research paper · 1 sourceEnergy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse TerrainsIn this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT).
- Oct 7, 2026 · Research paper · 1 sourceHow to train your model organismWe re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities.
- Oct 7, 2026 · Research paper · 1 sourceVisible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code GenerationVisible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior.
- Oct 7, 2026 · Research paper · 1 sourceCertified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration BudgetsWe ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces.
- Oct 6, 2026 · Research paper · 1 sourceEgoLAP: Learning from Egocentric Human Data through Language-Action ReasoningWe introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought.
- Oct 6, 2026 · Research paper · 1 sourcePOLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM AgentsWe propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology.
- Oct 6, 2026 · Research paper · 1 sourceThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning ModelsWe propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments.
- Oct 2, 2026 · Opinion / analysis · 1 sourceDon’t be fooled—LLMs don’t reasonAlphaGo won the game, ultimately triumphing 4-1 over Lee Sedol, one of the greatest professional Go players of all time. “I thought AlphaGo was based on probability calculation and that it was merely a machine,” Lee said afterwards. “But when I saw this move, I changed my mind.
- Oct 1, 2026 · Product / feature launch · 3 sources[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAUIn any case, you have any number of recaps coming at you today, and we’ll be shipping our DevDay pod soon, so you can either watch the full 1 hour livestream or this 15 minute supercut:
- Sep 30, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.18.0: Release 5.18.0Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
- Sep 29, 2026 · Tutorial / explainer · 1 sourceHow to Stop AI Agents From Secretly CollaboratingThe most famous example is OpenAI’s hack of AI platform Hugging Face, in which a swarm of roughly 700 AI agents escaped a testing environment and then hacked several companies, searching for information that could help them disguise cheating on a cybersecurity benchmark called ExploitGym.
- Sep 29, 2026 · Opinion / analysis · 1 sourceThe Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language ModelsThe Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
- Jun 16, 2026 · Opinion / analysis · 1 sourceUnlocking UK house-building with AI-accelerated planningNew UK government AI planning prototype built with Gemini aims to halve the time it takes to process homeowner applications
- Jun 16, 2026 · Opinion / analysis · 1 sourceSecuring the future of AI agentsHow we’re securing internal systems against increasingly capable and imperfectly aligned AI