Research
Papers and datasets worth knowing, ranked by significance and community attention.
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation.
Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.
BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech
While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored.
Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR
We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).
No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation
Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.
MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair
We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data.
Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.
Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening
MemoCare is an interactive mobile system for automated multimodal cognitive screening.
Steerspeech: Activation Steering For Emotion Control In Generated Speech
We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed.
Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation
We introduce cylindrical geodesic flow matching for paired cardiovascular waveform translation.
Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals
Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance.
Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR
This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM).
DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.
Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.
BeatFlow-ECG: Rectified Flow for ECG Reconstruction from Indirect Wearable Signals
We present BeatFlow-ECG, a conditional rectified-flow model for reconstructing single-channel ECG from synchronized PPG and inertial measurements.
From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis
We define a customer context graph as a unified model of customer and business context.
TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning
We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents.
MS-ECG-FM: Towards a More Universal Electrocardiogram Foundation Model for Health Monitoring using Multi-source Contrastive Learning
We introduce a new ECG foundation model --- MS-ECG-FM --- that is trained through contrastive alignment to multiple distinct clinical note types, including ECG, echocardiography, radiology, and discharge reports.
VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation
Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process.
Backdooring Acoustic Foundation Models for Physically Realizable Triggers
Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated.
BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
A Strength-Monotonic Law for Domain Alignment in Frozen-Embedding Bioacoustic Classification
When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains?
A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers
In a meta-analysis of published work, we show that automated judgement is the present norm.
Breaking Adversarial Transferability in Fine-Tuned Speech Recognition
We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients.
MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference.
Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools
Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements.
PVSync: A Unified Lip-Sync Expert for Timing and Articulation
We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring.
HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research.
How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders
Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec).
DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion
Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame.
Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization
Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection.
A persistent accuracy ceiling in automated verbal deception detection
Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines.
Could LLM Watermark Detection be Public?
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback.
Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents
To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory.
Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.
Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy
Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function.
Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.