Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Paper
Hugging Face Daily Papers2 sources3d ago

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.

▲ 37 upvotesPaper
Video
AI Engineer (YouTube)7h ago

Teaching Agents to Search with NVIDIA Data Designer — Dhruv Nathawani, NVIDIA

Dhruv Nathawani uses that funnel to explain how NVIDIA builds synthetic data that teaches a model to search, rather than answer from memory.

WIRED: AI12h ago

I Made Terrible Games With Google’s AI Playground

A long day’s haul in the video game slop mines.

OpenAI News3d ago

Pollo AI turns creative ideas into campaigns with OpenAI

With GPT-5.6, GPT-6 Astra, and GPT‐Image‐2.5, Pollo AI helps creators turn bold ideas into detailed images and cinematic video ads.

Paper
Apple Machine Learning Research2 sources5d ago

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.

Paper
Microsoft Research Blog4d ago

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.

Vendor claim only
Hacker News: LLM threads (40+ points)2 sources3d ago

Show HN: Jevman – AI decision models play Pac-Man

Openai just launched their decisions endpoint, cloudflare launched clef the other week, and many more jev alternatives are out there.

GitHub: unslothai/unsloth3d ago

unslothai/unsloth v0.1.905-beta: Sandboxing is here!

We're introducing Windows, Mac and Linux sandboxing in Unsloth!

Code
GitHub: anthropics/claude-code3d ago

anthropics/claude-code v2.1.294

Fixed prompt and agent hooks written as instructions (such as "Block commands that...") allowing what they should block

Code
MarkTechPost2d ago

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on.

GitHub: huggingface/trl5d ago

huggingface/trl v1.14.2

Patch release fixing two cases of silently wrong training and three crashes.

Code
Latent Space3d ago

Synthesis Superintelligence: from Semiconductors to Superconductors — Periodic Labs’ Liam Fedus and Ekin Dogus Cubuk

We go deep on Periodic’s vision for “synthesis superintelligence”: reinforcement learning grounded in physical experiments, AI-powered materials characterization, simulations and density functional theory, high-throughput labs, and systems that learn from the entire process of doing science rather than only its published results.

Hugging Face Blog4d ago

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.

Vendor claim only
Cloudflare Blog: AI10d ago

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required.

Anthropic Research12d ago

GLM-5.3 and the spread of advanced cyber capabilities

Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.

GitHub: pydantic/pydantic-ai9d ago

pydantic/pydantic-ai v2.53.0: v2.53.0 (2026-10-01)

This release fixes one security issue in ConcurrencyLimitedModel.

Code
Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Q-Learning with Scalar Adjoint Matching

Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.

Paper
Video
AI Engineer (YouTube)4h ago

Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI

Open-weight models are good enough now.

Paper
Hugging Face Daily Papers2 sources4d ago

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.

Paper
Video
NewAI Engineer (YouTube)2h ago

Hill-Climbing Skills: Improve Agents Without Changing the Model — Shubhankar Srivastava, Browserbase

Shubhankar Srivastava uses that uneven progress to show how browser agents can learn a task without changing model weights.

Paper
Hugging Face Daily Papers2 sources5d ago

Agent Plasticity: Measuring Self-Improvement Through Experience

We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance.

Paper
Paper
Hugging Face Daily Papers3d ago

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

System Switch: When Should a Fast Decision Model Stop and Think?

Dual-process agents pair a fast policy with a slow deliberative model.

Paper
OpenAI News4d ago

Radisson Hotel Group brings hotel discovery into ChatGPT

Radisson partnered with Accenture to build a ChatGPT plugin using OpenAI technology, helping travelers find, compare, and book hotels while planning their trips.

Vendor claim only
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution

Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards

We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.

Paper
OpenAI News5d ago

How Jump Trading is scaling quant research with ChatGPT

Jump Trading uses OpenAI to expand quantitative research.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Safe Meta-Policy Design with Risk Control

We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition

Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging.

Paper
GitHub: unslothai/unsloth4d ago

unslothai/unsloth v0.1.904-beta: Train your own Decision model

Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.

Code
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata

In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev

In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cost-Aware Mixture-of-Experts Coordination for Model Markets

This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Executing Causal Structure Learning with Linear-Attention Transformers

We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity.

Paper
Paper
Hugging Face Daily Papers5d ago

A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning

We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Shared and structured inputs undermine collective random choice by reasoning AI agents

Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Verification and Self-Improvement in Agentic AI: Foundations and Limits

Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs.

Paper