Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Paper
OpenAI News10 sources1d ago

Sharing AI progress in mathematics

OpenAI publishes new results on open problems in mathematics from an internal frontier model and shares Lean proof formalizations and research details on GitHub.

8 outlets
Understanding AI (Timothy B. Lee)3 sources2d ago

Understanding Jev, the new model everyone is talking about

On September 15, the startup TypeSafe AI came out of stealth and released a new AI model.

3 outlets
Paper
Hugging Face Daily Papers2 sources3d ago

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.

▲ 34 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)2 sources1d ago

Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review

Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task.

Paper
Hugging Face trending models7d ago

nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2

nerkyor published the model Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 on Hugging Face.

Weights
MarkTechPost15h ago

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context

OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.

Google DeepMind Blog4 sources10d ago

Gemini 4 Argon: our next era of frontier intelligence

Gemini 4 Argon: our next era of frontier intelligence

3 outlets
Import AI (Jack Clark)2 sources6d ago

Why agent swarms could be the next “scaling law”

One of the most surprising aspects of July’s news that OpenAI agents attacked Hugging Face was how the agents had worked together.

2 outlets
Guide to AI (Nathan Benaich)3d ago

The State of AI Report 2026

After months of research and revisions right up to the last minute, I’m thrilled to bring you the 9th annual State of AI Report.

Simon Willison's Weblog6d ago

Qwen3.8 27B addition in words

Research: Qwen3.8 27B addition in words

Ben's Bites12d ago

Sonnet 5.5 is worth a try

I’ll be at OpenAI DevDay today.

The Kaitchup (Benjamin Marie)5d ago

Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh

Like Qwen3.8 27B, Qwen3.8 Flash Next has three reasoning efforts: low, medium, and xhigh.

Latent Space9d ago

Academia is for Ambition — Alex Zhang, MIT

Last call for regular tickets for AI Engineer NYC!

Video
Machine Learning Street Talk (YouTube)11d ago

The Programming Language That Referees Mathematics — Leo de Moura

Leonardo de Moura created Lean and co-created Z3.

Show HN: AI projects (15+ points)11d ago

Show HN: Lathoa, a math app for kids where the AI is wrong on purpose

A robot called Errol solves a math problem step by step and one of the steps is wrong.

Paper
Hugging Face Daily Papers2 sources5d ago

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models

Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals

In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

On-Policy Distillation Teaches New Skills but Not New Knowledge

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Large language models are informing decisions with ever-higher stakes.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

BrickBench: Evaluating Agentic Brick Design

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

We propose SanSi, which turns a pre-trained looped language model into a typed decision model.

Paper
Paper
Hugging Face Daily Papers3d ago

TestPrism: Rethinking Test Evaluation Beyond a Single Reference

Large language model (LLM) coding agents have advanced test generation across diverse programming tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning

We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers

We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming

Developing optimization models for production scheduling requires substantial expert effort.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks

We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

We study how a coding agent learns across a sequence of abstract reasoning tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Code Understanding is a Bottleneck for Coding Agents

We present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

On the estimation and validity of AI time horizons---a statistical look at the METR plot

On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation

Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation

PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Relational Abstractions for Spatial Reasoning with Diffusion Models

To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment

We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence.

Paper
OpenAI News4d ago

Helping teens learn, plan, and shape the future of AI

College Planner is coming to ChatGPT for Teens to help students manage college applications, alongside new flashcards, quizzes, and a teen AI council.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry

We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Sparse Planning in Visual World Models via Cost Gradients

We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token.

Paper
OpenAI News9d ago

A model guide for the GPT-6 family

Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.