Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources3d ago

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.

▲ 52 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings

To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.

▲ 37 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.

▲ 24 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

▲ 71 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Reasoning-Informed Visual Editing

To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.

▲ 21 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video

We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.

▲ 24 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

SuperNav: An Agentic Navigation System for Any Task in Any Scene

General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.

▲ 71 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

VibeEdit: Image Editing with Canvas Instructions

We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.

▲ 18 upvotesPaper
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models

Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

V-CoLA: Vision Token Compression with Linear Attention

To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

SPW-Nav Streams Language-Guided Panoramic Video in Real Time

Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DataVista: Diagnosing Multimodal LLMs on Data Video Understanding

We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation

Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

On the Necessity of Attention-FFN Split in Vision Transformers

In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute

We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).

Paper
Paper
Hugging Face Daily Papers3d ago

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers

Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding

Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

4-Tensor Attention Model for Semantic Physical Reality

We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception

To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation

We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models

Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams

We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MultiWorldBench: Do Independently Controlled Views Describe One Shared World?

We introduce MultiWorldBench, a diagnostic Minecraft benchmark containing 495 case configurations across seven task suites and ten capabilities, including independent control, cross-view motion, shared-state synchronization, persistence, structural reasoning, concurrent interaction, and delayed revisit.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Seek-and-View Reasoning for Multi-View Spatial Understanding

To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization

To address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding

We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding

Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders

Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs

While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams.

Paper