Technique

Chain-of-thought

Also known as: CoT, chain of thought

19stories this week
24last 30 days
26all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Every Model That Can Be Run On 10-16GB VRAM Ranked
    It's been 4 years since c.ai first hallucination model, and yet, we're nowhere good enough at LLMs in terms of spontaneity/interesting hallucination features.
  2. Oct 8, 2026 · Research paper · 1 source
    Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
    Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.
  3. Oct 8, 2026 · Research paper · 1 source
    From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
    We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.
  4. Oct 8, 2026 · Research paper · 1 source
    Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
    Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
  5. Oct 8, 2026 · Research paper · 2 sources
    SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
    Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
  6. Oct 8, 2026 · Research paper · 1 source
    Deception by Omission: Language Models Knowingly Hide Their Mistakes
    Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed.
  7. Oct 8, 2026 · Research paper · 1 source
    Lapras: Latent Reasoning for Time Series Language Models
    Time Series Language Models (TSLMs) offer a promising path toward time series understanding by reasoning over temporal signals and producing natural language answers and explanations.
  8. Oct 7, 2026 · Research paper · 1 source
    SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
    Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
  9. Oct 7, 2026 · Research paper · 1 source
    KDFP: A first-principles approach to knowledge distillation in large language models
    Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models.
  10. Oct 7, 2026 · Research paper · 1 source
    EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
    We propose EngramEdit for decoupled knowledge updates through conditional memory.
  11. Oct 7, 2026 · Research paper · 1 source
    Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
    Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.
  12. Oct 7, 2026 · Research paper · 1 source
    Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
    Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
  13. Oct 7, 2026 · Research paper · 1 source
    Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains
    In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT).
  14. Oct 7, 2026 · Research paper · 1 source
    How to train your model organism
    We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities.
  15. Oct 7, 2026 · Research paper · 1 source
    Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation
    Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior.
  16. Oct 7, 2026 · Research paper · 1 source
    Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets
    We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces.
  17. Oct 6, 2026 · Research paper · 1 source
    EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
    We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought.
  18. Oct 6, 2026 · Research paper · 1 source
    POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
    We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology.
  19. Oct 6, 2026 · Research paper · 1 source
    ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models
    We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments.
  20. Oct 2, 2026 · Opinion / analysis · 1 source
    Don’t be fooled—LLMs don’t reason
    AlphaGo won the game, ultimately triumphing 4-1 over Lee Sedol, one of the greatest professional Go players of all time. “I thought AlphaGo was based on probability calculation and that it was merely a machine,” Lee said afterwards. “But when I saw this move, I changed my mind.
  21. Oct 1, 2026 · Product / feature launch · 3 sources
    [AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU
    In any case, you have any number of recaps coming at you today, and we’ll be shipping our DevDay pod soon, so you can either watch the full 1 hour livestream or this 15 minute supercut:
  22. Sep 30, 2026 · Open-source release · 1 source
    huggingface/transformers v5.18.0: Release 5.18.0
    Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
  23. Sep 29, 2026 · Tutorial / explainer · 1 source
    How to Stop AI Agents From Secretly Collaborating
    The most famous example is OpenAI’s hack of AI platform Hugging Face, in which a swarm of roughly 700 AI agents escaped a testing environment and then hacked several companies, searching for information that could help them disguise cheating on a cybersecurity benchmark called ExploitGym.
  24. Sep 29, 2026 · Opinion / analysis · 1 source
    The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
    The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
  25. Jun 16, 2026 · Opinion / analysis · 1 source
    Unlocking UK house-building with AI-accelerated planning
    New UK government AI planning prototype built with Gemini aims to halve the time it takes to process homeowner applications
  26. Jun 16, 2026 · Opinion / analysis · 1 source
    Securing the future of AI agents
    How we’re securing internal systems against increasingly capable and imperfectly aligned AI

Often appears with