Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

▲ 71 upvotesPaper
Video
NewAI Engineer (YouTube)1h ago

Stop Renting Intelligence: The Train-to-Deploy Loop for Specialized AI — Fireworks AI

Jetashree Ravi, who leads part of the applied machine learning team at Fireworks AI, explains how teams move from closed models to open ones without losing quality.

Latent Space1d ago

Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub

From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology.

r/LocalLLaMA (top, daily)18h ago

Converting dense models into Mixture-of-Experts

For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.

Paper
Microsoft Research Blog4d ago

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.

Vendor claim only
Google Research Blog9d ago

Toward provably private learning from federated data

In 2017, Google introduced Federated Learning (FL) a machine learning technique that trains models across decentralized, private data.

Vendor claim only
r/MachineLearning (top, daily)9h ago

I trained a 414k-parameter transformer to fly a boids flock, then tested whether the rules a probe can read are the ones it uses [P]

I wrote a small boid simulator (12 birds), recorded it flying, and trained a transformer to predict each bird's next move without it knowing about any boid rules.

NVIDIA Technical Blog5d ago

Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly.

Vendor claim only
MarkTechPost2d ago

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on.

r/StableDiffusion (top, day)1d ago

Krea2 Turbo Distill 2 step LoRA - FINAL checkpoint released (chk51195)

Krea 2 Turbo — 2-Step Distillation LoRA (FINAL Version)

Interconnects (Nathan Lambert)2d ago

I expect rapid progress but not towards general superintelligence

I’ve often been surprised when I hear from top researchers in industry that they think AI will be better than them at their job in a few years, and I didn’t really know why I doubted it.

Hugging Face Blog4d ago

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.

Vendor claim only
GitHub: unslothai/unsloth10d ago

unslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UX

This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.

Code
GitHub: pytorch/pytorch11d ago

pytorch/pytorch v2.14.1: PyTorch 2.14.1 Release

This release is meant to fix the following regressions and silent correctness issues:

Code
Epoch AI: Gradient Updates13d ago

AI is getting cheaper faster than any other transformative technology

Because that’s how fast AI is getting cheaper.

GitHub: huggingface/peft12d ago

huggingface/peft v0.21.1

This is a small PEFT release to enable Tensor Parallel (TP) to work properly with PEFT.

Code
Deep (Learning) Focus (Cameron Wolfe)13d ago

Notes on NVIDIA Nemotron

Today, many key details of frontier large language models (LLMs) remain proprietary, but open-weights model families—such as DeepSeek, Kimi, and MiMo—continue to provide a valuable window into the development process for modern LLMs. Among these resources, the NVIDIA Nemotron model series is especially useful due to its transparency.

Apple Machine Learning Research11d ago

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

Paper
Hugging Face Daily Papers2 sources3d ago

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models

Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.

Paper
Video
NewAI Engineer (YouTube)2h ago

Hill-Climbing Skills: Improve Agents Without Changing the Model — Shubhankar Srivastava, Browserbase

Shubhankar Srivastava uses that uneven progress to show how browser agents can learn a task without changing model weights.

Video
NewAI Engineer (YouTube)2h ago

Parameter Golf with AutoResearch — Vayum Arora, Zhengyao Jiang, Dixing Xu & Dhruv Srikanth, Weco AI

Zhengyao Jiang introduces autoresearch as repeated proposals and evaluations, and Dixing Xu explains the team's Aiden system and its contributions to OpenAI's Parameter Golf challenge.

Paper
Hugging Face Daily Papers2 sources3d ago

Predicting Cable Dynamics with Physical Attention Bias

Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution

Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MotherTree: Meta-learning on synthetic data improves decision tree training

We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

EDiS: Edge Disjoint Subgraph Sparsification Framework for Graph Neural Networks

We introduce EDiS (Edge-Disjoint Subgraph sparsification framework), which separates one-time structural extraction from per-epoch graph composition.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata

In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

asdex: Automatic Sparse Differentiation in JAX

Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cost-Aware Mixture-of-Experts Coordination for Model Markets

This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.

Paper
NVIDIA Technical Blog11d ago

Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK

AI infrastructure engineers, storage developers, and cloud service providers need fast and secure access to high-capacity file and object storage to support AI...AI infrastructure engineers, storage developers, and cloud service providers need fast and secure access to high-capacity file and object storage to support AI workloads.

Vendor claim only
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Executing Causal Structure Learning with Linear-Attention Transformers

We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Leaner Transformers Can Easily Learn to Cluster

Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising

We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Minimax Gaussian Mechanisms for Continual Machine Unlearning

We develop Gaussian mechanisms for Newton updates under sequential deletion requests.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction

We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction

Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Exact-Solution Volume and Length Generalization in Transformers

Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solution, if learned, is generalizable to longer input lengths.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples

Adversarial distillation transfers robustness from high-capacity teachers to compact students.

Paper