Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Paper
OpenAI News10 sources1d ago

Sharing AI progress in mathematics

OpenAI publishes new results on open problems in mathematics from an internal frontier model and shares Lean proof formalizations and research details on GitHub.

8 outlets
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)2 sources22h ago

Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review

Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task.

Paper
MarkTechPost14h ago

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context

OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.

Understanding AI (Timothy B. Lee)3 sources1d ago

Understanding Jev, the new model everyone is talking about

On September 15, the startup TypeSafe AI came out of stealth and released a new AI model.

3 outlets
Paper
Hugging Face Daily Papers2 sources3d ago

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.

▲ 34 upvotesPaper
Guide to AI (Nathan Benaich)3d ago

The State of AI Report 2026

After months of research and revisions right up to the last minute, I’m thrilled to bring you the 9th annual State of AI Report.

Import AI (Jack Clark)2 sources6d ago

Why agent swarms could be the next “scaling law”

One of the most surprising aspects of July’s news that OpenAI agents attacked Hugging Face was how the agents had worked together.

2 outlets
Simon Willison's Weblog2d ago

Quoting Carson Gross

Computer programming is, fundamentally, about two things:

The Kaitchup (Benjamin Marie)5d ago

Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh

Like Qwen3.8 27B, Qwen3.8 Flash Next has three reasoning efforts: low, medium, and xhigh.

Paper
Hugging Face Daily Papers2 sources3d ago

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals

In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

BrickBench: Evaluating Agentic Brick Design

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models

Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

On-Policy Distillation Teaches New Skills but Not New Knowledge

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers

We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.

Paper
Paper
Hugging Face Daily Papers3d ago

TestPrism: Rethinking Test Evaluation Beyond a Single Reference

Large language model (LLM) coding agents have advanced test generation across diverse programming tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

On the estimation and validity of AI time horizons---a statistical look at the METR plot

On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Large language models are informing decisions with ever-higher stakes.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

We study how a coding agent learns across a sequence of abstract reasoning tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation

PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving

Multimodal large language models (MLLMs) often struggle to identify and use geometric relations in diagrams.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry

We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Closed-loop evaluation of LLM agents for embedded software development

We present a benchmark for closed-loop evaluation of embedded coding agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance

Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small-world connectivity as a structural signature of LLM reasoning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation

We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming

Developing optimization models for production scheduling requires substantial expert effort.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL

A verified reference solution provides a correct trajectory for training a reasoning model.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Chronos Enables Code Agents to Reason over Software Evolution

We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation

Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation

Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning

We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Sparse Planning in Visual World Models via Cost Gradients

We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent Systems

We propose Error-Propagation Modeling for Failure Attribution (EMFA).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs

Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood.

Paper
OpenAI News4d ago

Helping teens learn, plan, and shape the future of AI

College Planner is coming to ChatGPT for Teens to help students manage college applications, alongside new flashcards, quizzes, and a teen AI council.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Relational Abstractions for Spatial Reasoning with Diffusion Models

To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Code Understanding is a Bottleneck for Coding Agents

We present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluation.

Paper