Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image.
VideoCan Your Agent Hear You Now? Building Live Voice Agents with Gemini — Thor Schaeff
An agent that hears you, sees you, and answers in your language, in real time.

These execs think voice AI hasn’t reached its ChatGPT moment yet
Voice AI's often misses important points for its context layer, and causes the whole pipeline to break
mistralai/Voxtral-Mini-4B-Realtime-Arabic
mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
pydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)
<!-- Release notes generated using configuration in .github/release.yml at main -->
A new feature for my blog, built using my voice
I used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment.
open-webui/open-webui v0.12.0
Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.
Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files.
Claude Dashboards & Motion
Ask Claude for live dashboards and animated explainers
Introducing Falcon ASR
We’re introducing Falcon-ASR, our 1.6 billion parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect.
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
Build a voice travel concierge with Amazon Bedrock AgentCore, Managed Knowledge Base and Nova Sonic
Airlines already have apps and websites where travelers check flights, pick seats, and manage bookings, and adding a natural voice layer opens those tasks to spoken requests.
unslothai/unsloth v0.1.903-beta: New Browser + Voice Cloning
This release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
Willow taps CoreWeave to simplify AI model training
Willow Care Inc., the company behind AI dictation app Willow Voice, chose CoreWeave Inc. for its forward-deployed support rather than compute alone.
mistralai/LIDstral-Arabic
mistralai published the model LIDstral-Arabic on Hugging Face.
MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.
Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.
Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech
While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored.
Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.
Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR
We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).
Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.
DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.
MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening
MemoCare is an interactive mobile system for automated multimodal cognitive screening.
Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
Steerspeech: Activation Steering For Emotion Control In Generated Speech
We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.
HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research.
How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders
Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec).
DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion
Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame.
Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization
Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection.
A persistent accuracy ceiling in automated verbal deception detection
Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines.
Could LLM Watermark Detection be Public?
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback.
SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation
Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.
Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents
To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory.
Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.
Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy
Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function.
BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.
Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition.
Prosody-to-Text: Predicting text from low-pass filtered speech
While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked.
When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing
We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing.
Backdooring Acoustic Foundation Models for Physically Realizable Triggers
Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated.