Research
Papers and datasets worth knowing, ranked by significance and community attention.
PaperSharing AI progress in mathematics
OpenAI publishes new results on open problems in mathematics from an internal frontier model and shares Lean proof formalizations and research details on GitHub.
PaperBeyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task.
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
PaperNormalizing Trajectory Models
We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory.
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.
Reasoning-Informed Visual Editing
To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
SuperNav: An Agentic Navigation System for Any Task in Any Scene
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide.
VibeEdit: Image Editing with Canvas Instructions
We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
AgentGarten: Code Worlds for Evolving Agents
We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments.
TokenRouter: Efficient Serving System for Token-Level LLM Routing
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving.
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
REMORY: Learning Residual Memory for Context Compaction
We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces.
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes.
V-CoLA: Vision Token Compression with Linear Attention
To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.
Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.
Pumpire: Unified Benchmark for Metric Distance Estimation
We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors.
Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments.
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework.
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance.
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed.
SpaceFlow: Locally Controllable 3D Generation
We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives.
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.
From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope.
BrickBench: Evaluating Agentic Brick Design
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.
Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
RoboJEPA: Scaling Laws for Multi-Embodiment Robotic Latent World Models
Researchers introduced RoboJEPA, an 8B-parameter multi-embodiment latent world model that establishes compute scaling laws and enables zero-shot real-robot planning toward goal images.
The Lattice of Transition Laws
Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens.
USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation
To address these limitations, we formulate articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence and introduce USDCraft, a framework in which a pretrained LLM writes and revises executable programs for simulation-ready articulated assets without task-specific training.
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.
Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.
Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
Predicting Cable Dynamics with Physical Attention Bias
Learned simulators for deformable linear objects (DLOs) such as cables have to predict the motion of cables they were not trained on and stay stable over long rollouts.