Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

r/LocalLLaMA (top, daily)15h ago

Converting dense models into Mixture-of-Experts

For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.

r/MachineLearning (top, daily)6h ago

I trained a 414k-parameter transformer to fly a boids flock, then tested whether the rules a probe can read are the ones it uses [P]

I wrote a small boid simulator (12 birds), recorded it flying, and trained a transformer to predict each bird's next move without it knowing about any boid rules.

Latent Space1d ago

Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub

From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology.

Paper
Hugging Face Daily Papers2 sources3d ago

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.

▲ 7 upvotesPaper
r/StableDiffusion (top, day)22h ago

Krea2 Turbo Distill 2 step LoRA - FINAL checkpoint released (chk51195)

Krea 2 Turbo — 2-Step Distillation LoRA (FINAL Version)

Interconnects (Nathan Lambert)1d ago

I expect rapid progress but not towards general superintelligence

I’ve often been surprised when I hear from top researchers in industry that they think AI will be better than them at their job in a few years, and I didn’t really know why I doubted it.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution

Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).

Paper
MarkTechPost2d ago

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on.

Microsoft Research Blog4d ago

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.

Vendor claim only
NVIDIA Technical Blog5d ago

Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly.

Vendor claim only
Hugging Face Blog4d ago

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.

Vendor claim only
Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models

Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Predicting Cable Dynamics with Physical Attention Bias

Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

asdex: Automatic Sparse Differentiation in JAX

Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cost-Aware Mixture-of-Experts Coordination for Model Markets

This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MotherTree: Meta-learning on synthetic data improves decision tree training

We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising

We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Minimax Gaussian Mechanisms for Continual Machine Unlearning

We develop Gaussian mechanisms for Newton updates under sequential deletion requests.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Recursive Self-Improvement through Multi-Agent Self-Supervision

To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction

Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution Drift

Accurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights

We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples

Adversarial distillation transfers robustness from high-capacity teachers to compact students.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DataSense-Bench: The First Step Toward an AI Scientist

We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport

To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Executing Causal Structure Learning with Linear-Attention Transformers

We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation

We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Generative Adversarial Loops

To enable self-advancing systems, we propose Generative Adversarial Loop (GAL), a generator-discriminator framework alternating between two agentic searches: (1) a discriminator that generates adversarial data to expose weaknesses in current systems, and (2) a generator that discovers algorithms to overcome them.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Just Weather Scoring: Efficient End-to-end Nowcasting with Distributional Diffusion

We introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Correlational Training of Morphological Neural Networks

In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Timer-M1: A Multivariate Time Series Foundation Model via Learning Primitives

We introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis

Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Agentic-TTT: Training test-time policy for test-time training

To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Learning Kilometer-Scale Weather Prediction with Global-Regional Alignment

We propose ScaleCast, a regional forecasting framework that addresses these challenges through Global-Regional Alignment.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance

We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models

With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Leaner Transformers Can Easily Learn to Cluster

Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Is Real-World Training Data Necessary for Generalist Graph Anomaly Detection?

Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning.

Paper