Research
Papers and datasets worth knowing, ranked by significance and community attention.
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
asdex: Automatic Sparse Differentiation in JAX
Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.
SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
TokenRouter: Efficient Serving System for Token-Level LLM Routing
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving.
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
V-CoLA: Vision Token Compression with Linear Attention
To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.
Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed.
Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
The Lattice of Transition Laws
Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.
Incremental Open-Ended Deep Research with Structured Harness
We introduce Incremental Open-Ended Deep Research (Incremental-OEDR), a research setting that treats a report as an evolving research state and incrementally updates it by preserving valid knowledge, revising outdated or incomplete content, and incorporating newly available information.
Opera: A Verbal Critic Framework for Long-horizon Coding Agents
We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved.
Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates.
From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.
Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
CARE: Certifying Acceleration for Vision-Language-Action Inference
Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.
Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.
Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising
We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization.
Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference.
Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding
Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch.
An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation
PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.
DSReg: Provably Recovering Individual World Latents without Reconstruction
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity.
La-Ribo: RNA Co-Design via Geometry-Latent Flow Matching
We introduce La-Ribo, a generative framework for RNA sequence-structure co-design via geometry-latent flow matching.
Traceable World State: A Provenance-Aware State Representation and Deterministic Replay Framework for Robotic Systems
We present Traceable World State (TWS), a middleware-neutral semantic representation and reference runtime for provenance-aware robot world state.
Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged.
Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
Compile the Table: Query-Calibrated Operator Compression for Tabular In-Context Learning
We propose QCOC (Query-Calibrated Operator Compression), which exploits the exchangeability and repeated use of in-context examples by compiling their full KV cache once into compact memory shared across subsequent queries.
RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
We present RFChipAgent, a first-of-its-kind multi-agent flow of large language model (LLM) agents for end-to-end analog/RF circuit design automation, in which AI agents collaboratively orchestrate the complete design flow under human supervision.
LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models
We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted.
Read What Matters: Query-Adaptive Quantization for KV Caches
KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation
We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.
When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.
RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
To bridge this gap, we introduce RaReCache, a framework that enables a large target model to decode accurately from a cache prefilled by a much smaller source via selective recomputation.
LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs
Lifelong editing of LLMs requires storing thousands of edits after acquisition.
ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights
We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.
Closed-loop evaluation of LLM agents for embedded software development
We present a benchmark for closed-loop evaluation of embedded coding agents.
$C_4$-Equivariant Flow Matching on Anisotropic Power-Diagram Graphs for Microstructure Generation
We introduce a generative model for synthesising realistic polycrystalline microstructures using flow matching and graph neural networks.
Executing Causal Structure Learning with Linear-Attention Transformers
We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.
ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count.
Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference
We introduce Finsler Flow Matching (FFM), a framework for learning continuous stochastic dynamics from discrete Markov transition graphs.
MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution
We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters.
OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature.
TR-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region Reformulation
Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers.