Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.

Paper
Paper
Hugging Face Daily Papers7d ago

SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision

We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation.

▲ 10 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Temporal transformer CAN encoder with federated lightweight heads for anomaly detection

Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech

While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR

We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation

Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification

We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair

We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders

Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SP-DocReader: Difference-Aware Self-Play for Precise Document OCR

We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening

MemoCare is an interactive mobile system for automated multimodal cognitive screening.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Steerspeech: Activation Steering For Emotion Control In Generated Speech

We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must

This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation

We introduce cylindrical geodesic flow matching for paired cardiovascular waveform translation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals

Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams

We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR

This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence

We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

BeatFlow-ECG: Rectified Flow for ECG Reconstruction from Indirect Wearable Signals

We present BeatFlow-ECG, a conditional rectified-flow model for reconstructing single-channel ECG from synchronized PPG and inertial measurements.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis

We define a customer context graph as a unified model of customer and business context.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MS-ECG-FM: Towards a More Universal Electrocardiogram Foundation Model for Health Monitoring using Multi-source Contrastive Learning

We introduce a new ECG foundation model --- MS-ECG-FM --- that is trained through contrastive alignment to multiple distinct clinical note types, including ECG, echocardiography, radiology, and discharge reports.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation

Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Backdooring Acoustic Foundation Models for Physically Realizable Triggers

Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics

In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

A Strength-Monotonic Law for Domain Alignment in Frozen-Embedding Bioacoustic Classification

When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains?

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers

In a meta-analysis of published work, we show that automated judgement is the present norm.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Breaking Adversarial Transferability in Fine-Tuned Speech Recognition

We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos

Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools

Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

PVSync: A Unified Lip-Sync Expert for Timing and Articulation

We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing

To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders

Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion

Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization

Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

A persistent accuracy ceiling in automated verbal deception detection

Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Could LLM Watermark Detection be Public?

Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents

To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion

Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy

Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.

Paper