Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.
PaperSharing AI progress in mathematics
OpenAI publishes new results on open problems in mathematics from an internal frontier model and shares Lean proof formalizations and research details on GitHub.

Understanding Jev, the new model everyone is talking about
On September 15, the startup TypeSafe AI came out of stealth and released a new AI model.
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.
PaperBeyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task.
nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
nerkyor published the model Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 on Hugging Face.

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context
OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.

Gemini 4 Argon: our next era of frontier intelligence
Gemini 4 Argon: our next era of frontier intelligence

Why agent swarms could be the next “scaling law”
One of the most surprising aspects of July’s news that OpenAI agents attacked Hugging Face was how the agents had worked together.

The State of AI Report 2026
After months of research and revisions right up to the last minute, I’m thrilled to bring you the 9th annual State of AI Report.
Qwen3.8 27B addition in words
Research: Qwen3.8 27B addition in words

Sonnet 5.5 is worth a try
I’ll be at OpenAI DevDay today.

Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh
Like Qwen3.8 27B, Qwen3.8 Flash Next has three reasoning efforts: low, medium, and xhigh.
Academia is for Ambition — Alex Zhang, MIT
Last call for regular tickets for AI Engineer NYC!
VideoThe Programming Language That Referees Mathematics — Leo de Moura
Leonardo de Moura created Lean and co-created Z3.
Show HN: Lathoa, a math app for kids where the AI is wrong on purpose
A robot called Errol solves a math problem step by step and one of the steps is wrong.
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models
Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.
Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.
On-Policy Distillation Teaches New Skills but Not New Knowledge
On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.
U-Space: Uncovering When and Why Uncertainty Arises in Language Models
Large language models are informing decisions with ever-higher stakes.
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.
BrickBench: Evaluating Agentic Brick Design
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.
SanSi: A Looped Typed Decision Model for System 1.5 Thinking
We propose SanSi, which turns a pre-trained looped language model into a typed decision model.
TestPrism: Rethinking Test Evaluation Beyond a Single Reference
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks.
The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum.
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.
From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.
Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming
Developing optimization models for production scheduling requires substantial expert effort.
GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer.
Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
We study how a coding agent learns across a sequence of abstract reasoning tasks.
SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.
Code Understanding is a Bottleneck for Coding Agents
We present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluation.
RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.
Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.
On the estimation and validity of AI time horizons---a statistical look at the METR plot
On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time.
Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation
Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior.
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation
PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.
Relational Abstractions for Spatial Reasoning with Diffusion Models
To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations.
BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment
We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence.
Helping teens learn, plan, and shape the future of AI
College Planner is coming to ChatGPT for Teens to help students manage college applications, alongside new flashcards, quizzes, and a teen AI council.
DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.
ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
Sparse Planning in Visual World Models via Cost Gradients
We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token.
A model guide for the GPT-6 family
Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.