Research
Papers and datasets worth knowing, ranked by significance and community attention.
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.
Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech
While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored.
Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.
Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR
We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).
Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.
DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.
MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening
MemoCare is an interactive mobile system for automated multimodal cognitive screening.
Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
Steerspeech: Activation Steering For Emotion Control In Generated Speech
We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.
HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research.
How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders
Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec).
DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion
Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame.
Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization
Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection.
A persistent accuracy ceiling in automated verbal deception detection
Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines.
Could LLM Watermark Detection be Public?
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback.
SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation
Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.
Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents
To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory.
Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.
Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy
Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function.
BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.
Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition.
Prosody-to-Text: Predicting text from low-pass filtered speech
While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked.
When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing
We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing.
Backdooring Acoustic Foundation Models for Physically Realizable Triggers
Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated.
Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair
We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data.
A Strength-Monotonic Law for Domain Alignment in Frozen-Embedding Bioacoustic Classification
When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains?
Phonological Interference in Multilingual Speech Models
We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language.
From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis
We define a customer context graph as a unified model of customer and business context.
A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers
In a meta-analysis of published work, we show that automated judgement is the present norm.
TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning
We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents.
Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary.
VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation
Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process.
Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation
We introduce cylindrical geodesic flow matching for paired cardiovascular waveform translation.
Learning What to Trust in Multimodal Learning under Noisy Supervision
Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection.
Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals
Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance.
Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators
In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling.
BeatFlow-ECG: Rectified Flow for ECG Reconstruction from Indirect Wearable Signals
We present BeatFlow-ECG, a conditional rectified-flow model for reconstructing single-channel ECG from synchronized PPG and inertial measurements.
The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed.