Research
Papers and datasets worth knowing, ranked by significance and community attention.
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models
Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
Predicting Cable Dynamics with Physical Attention Bias
Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.
asdex: Automatic Sparse Differentiation in JAX
Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.
Cost-Aware Mixture-of-Experts Coordination for Model Markets
This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.
MotherTree: Meta-learning on synthetic data improves decision tree training
We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass.
Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.
Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.
Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising
We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization.
Minimax Gaussian Mechanisms for Continual Machine Unlearning
We develop Gaussian mechanisms for Newton updates under sequential deletion requests.
Recursive Self-Improvement through Multi-Agent Self-Supervision
To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories.
TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction
Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.
AdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution Drift
Accurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly.
Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights
We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.
Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples
Adversarial distillation transfers robustness from high-capacity teachers to compact students.
When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?
DataSense-Bench: The First Step Toward an AI Scientist
We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights.
OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step.
Executing Causal Structure Learning with Linear-Attention Transformers
We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.
Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood.
Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation
We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy.
Generative Adversarial Loops
To enable self-advancing systems, we propose Generative Adversarial Loop (GAL), a generator-discriminator framework alternating between two agentic searches: (1) a discriminator that generates adversarial data to expose weaknesses in current systems, and (2) a generator that discovers algorithms to overcome them.
Just Weather Scoring: Efficient End-to-end Nowcasting with Distributional Diffusion
We introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation.
Correlational Training of Morphological Neural Networks
In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme.
Timer-M1: A Multivariate Time Series Foundation Model via Learning Primitives
We introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting.
Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis
Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.
$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.
Agentic-TTT: Training test-time policy for test-time training
To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions.
Learning Kilometer-Scale Weather Prediction with Global-Regional Alignment
We propose ScaleCast, a regional forecasting framework that addresses these challenges through Global-Regional Alignment.
AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance
We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling.
Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models
With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
Leaner Transformers Can Easily Learn to Cluster
Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$).
Is Real-World Training Data Necessary for Generalist Graph Anomaly Detection?
Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning.
Connected Self Forcing: Beyond Local Learning in Video Autoregression
To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching.
Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains
To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets.
EDiS: Edge Disjoint Subgraph Sparsification Framework for Graph Neural Networks
We introduce EDiS (Edge-Disjoint Subgraph sparsification framework), which separates one-time structural extraction from per-epoch graph composition.
PulseBound: Future-Beat State Forecasting Under an Explicit Information Boundary
We introduce PulseBound, a PPG representation learner combining physiologically structured future-beat prediction with an explicit stored-window information boundary.
New Lower Bound and Upper Bounds on the Regret for Online Sparse Linear Regression
We study online sparse linear regression (OSLR) where any algorithm is restricted to accessing only $b$ out of $d$ attributes per instance for prediction and $b0\geq 0$ additional attributes after prediction, which was proved to be NP-hard.
Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs
GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference.
Teaching PPG How not Who: Fixed-Effects Distillation from ECG
ECG is widely used to teach PPG-only models, yet what it teaches is unexamined.
HAN-Mamba: Hierarchical Selective State Space Networks for Multi-Scale Financial Volatility Forecasting
Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift.
Velocity Scaling in Flow Matching
Scaling a learned flow-matching velocity field $vθ$ by a gain $γ(t)$ was recently shown to greatly improve generation quality.