Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
VideoTeaching Agents to Search with NVIDIA Data Designer — Dhruv Nathawani, NVIDIA
Dhruv Nathawani uses that funnel to explain how NVIDIA builds synthetic data that teaches a model to search, rather than answer from memory.
I Made Terrible Games With Google’s AI Playground
A long day’s haul in the video game slop mines.
Pollo AI turns creative ideas into campaigns with OpenAI
With GPT-5.6, GPT-6 Astra, and GPT‐Image‐2.5, Pollo AI helps creators turn bold ideas into detailed images and cinematic video ads.
PaperRISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.
Show HN: Jevman – AI decision models play Pac-Man
Openai just launched their decisions endpoint, cloudflare launched clef the other week, and many more jev alternatives are out there.
unslothai/unsloth v0.1.905-beta: Sandboxing is here!
We're introducing Windows, Mac and Linux sandboxing in Unsloth!
anthropics/claude-code v2.1.294
Fixed prompt and agent hooks written as instructions (such as "Block commands that...") allowing what they should block
Google Research RRSI Guide: Mastering Self-Improving AI Agents
In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on.
huggingface/trl v1.14.2
Patch release fixing two cases of silently wrong training and three crashes.
Synthesis Superintelligence: from Semiconductors to Superconductors — Periodic Labs’ Liam Fedus and Ekin Dogus Cubuk
We go deep on Periodic’s vision for “synthesis superintelligence”: reinforcement learning grounded in physical experiments, AI-powered materials characterization, simulations and density functional theory, high-throughput labs, and systems that learn from the entire process of doing science rather than only its published results.
One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform
While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required.

GLM-5.3 and the spread of advanced cyber capabilities
Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.
pydantic/pydantic-ai v2.53.0: v2.53.0 (2026-10-01)
This release fixes one security issue in ConcurrencyLimitedModel.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments.
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed.
ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
VideoSame Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now.
UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.
VideoHill-Climbing Skills: Improve Agents Without Changing the Model — Shubhankar Srivastava, Browserbase
Shubhankar Srivastava uses that uneven progress to show how browser agents can learn a task without changing model weights.
Agent Plasticity: Measuring Self-Improvement Through Experience
We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance.
Opera: A Verbal Critic Framework for Long-horizon Coding Agents
We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
System Switch: When Should a Fast Decision Model Stop and Think?
Dual-process agents pair a fast policy with a slow deliberative model.
Radisson Hotel Group brings hotel discovery into ChatGPT
Radisson partnered with Accenture to build a ChatGPT plugin using OpenAI technology, helping travelers find, compare, and book hotels while planning their trips.
Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
How Jump Trading is scaling quant research with ChatGPT
Jump Trading uses OpenAI to expand quantitative research.
Safe Meta-Policy Design with Risk Control
We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression.
On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.
Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging.
unslothai/unsloth v0.1.904-beta: Train your own Decision model
Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata
In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.
From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev
In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots.
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.
Cost-Aware Mixture-of-Experts Coordination for Model Markets
This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.
Executing Causal Structure Learning with Linear-Attention Transformers
We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.
PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.
A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.
Shared and structured inputs undermine collective random choice by reasoning AI agents
Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.
Verification and Self-Improvement in Agentic AI: Foundations and Limits
Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs.