Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

▲ 71 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.

▲ 7 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

▲ 8 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.

▲ 8 upvotesPaper
Paper
Hugging Face Daily Papers2 sources4d ago

RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models

Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.

▲ 5 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Predicting Cable Dynamics with Physical Attention Bias

Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.

▲ 5 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution

Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MotherTree: Meta-learning on synthetic data improves decision tree training

We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

EDiS: Edge Disjoint Subgraph Sparsification Framework for Graph Neural Networks

We introduce EDiS (Edge-Disjoint Subgraph sparsification framework), which separates one-time structural extraction from per-epoch graph composition.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata

In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

asdex: Automatic Sparse Differentiation in JAX

Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cost-Aware Mixture-of-Experts Coordination for Model Markets

This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Executing Causal Structure Learning with Linear-Attention Transformers

We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Leaner Transformers Can Easily Learn to Cluster

Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising

We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Minimax Gaussian Mechanisms for Continual Machine Unlearning

We develop Gaussian mechanisms for Newton updates under sequential deletion requests.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction

We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction

Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Exact-Solution Volume and Length Generalization in Transformers

Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solution, if learned, is generalizable to longer input lengths.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples

Adversarial distillation transfers robustness from high-capacity teachers to compact students.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models

With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Learning Transition Kernels of Jump-Diffusion Processes with Conditional Diffusion Models

We study the problem of learning transition kernels for time-homogeneous jump-diffusion processes using conditional diffusion models, with the goal of generating new sample paths from training data consisting of N independent trajectories observed on a high-frequency discrete time grid.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction

To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM$^2$ES) framework for survival prediction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Steering Diffusion Models to Rare Events with Sequential Monte Carlo

In this paper, we present Diffusion Importance Sampling of Rare Events or DireSMC, a sequential Monte Carlo scheme that guides a population of weighted samples towards the rare event, giving access not only to samples but also to a calibrated estimate of its probability.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Recursive Self-Improvement through Multi-Agent Self-Supervision

To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Teaching PPG How not Who: Fixed-Effects Distillation from ECG

ECG is widely used to teach PPG-only models, yet what it teaches is unexamined.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

HAN-Mamba: Hierarchical Selective State Space Networks for Multi-Scale Financial Volatility Forecasting

Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights

We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Multi-Agent Coordination via Support-Preserving Distillation

To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training

Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

AdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution Drift

Accurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation

We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Generative Adversarial Loops

To enable self-advancing systems, we propose Generative Adversarial Loop (GAL), a generator-discriminator framework alternating between two agentic searches: (1) a discriminator that generates adversarial data to expose weaknesses in current systems, and (2) a generator that discovers algorithms to overcome them.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

SAPD: Step-Aligned Privileged Distillation

We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model.

Paper