Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Cloudflare Blog: AI2 sources2d ago

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image.

Weights
Video
AI Engineer (YouTube)6h ago

Can Your Agent Hear You Now? Building Live Voice Agents with Gemini — Thor Schaeff

An agent that hears you, sees you, and answers in your language, in real time.

TechCrunch: AI7h ago

These execs think voice AI hasn’t reached its ChatGPT moment yet

Voice AI's often misses important points for its context layer, and causes the whole pipeline to break

HF: Mistral AI3d ago

mistralai/Voxtral-Mini-4B-Realtime-Arabic

mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.

Weights
GitHub: pydantic/pydantic-ai2d ago

pydantic/pydantic-ai v2.55.0: v2.55.0 (2026-10-09)

<!-- Release notes generated using configuration in .github/release.yml at main -->

Code
Simon Willison's Weblog2d ago

A new feature for my blog, built using my voice

I used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment.

GitHub: open-webui/open-webui1d ago

open-webui/open-webui v0.12.0

Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.

Code
r/LocalLLaMA (top, daily)1d ago

Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them

DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files.

Product Hunt: AI launches2d ago

Claude Dashboards & Motion

Ask Claude for live dashboards and animated explainers

Hugging Face Blog4d ago

Introducing Falcon ASR

We’re introducing Falcon-ASR, our 1.6 billion parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect.

Vendor claim only
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.

Paper
AWS Machine Learning Blog5d ago

Build a voice travel concierge with Amazon Bedrock AgentCore, Managed Knowledge Base and Nova Sonic

Airlines already have apps and websites where travelers check flights, pick seats, and manage bookings, and adding a natural voice layer opens those tasks to spoken requests.

GitHub: unslothai/unsloth5d ago

unslothai/unsloth v0.1.903-beta: New Browser + Voice Cloning

This release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.

Code
SiliconANGLE: AI3d ago

Willow taps CoreWeave to simplify AI model training

Willow Care Inc., the company behind AI dictation app Willow Voice, chose CoreWeave Inc. for its forward-deployed support rather than compute alone.

HF: Mistral AI3d ago

mistralai/LIDstral-Arabic

mistralai published the model LIDstral-Arabic on Hugging Face.

Weights
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification

We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must

This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams

We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SP-DocReader: Difference-Aware Self-Play for Precise Document OCR

We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Temporal transformer CAN encoder with federated lightweight heads for anomaly detection

Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech

While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders

Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence

We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR

We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening

MemoCare is an interactive mobile system for automated multimodal cognitive screening.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Steerspeech: Activation Steering For Emotion Control In Generated Speech

We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing

To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders

Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion

Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization

Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

A persistent accuracy ceiling in automated verbal deception detection

Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Could LLM Watermark Detection be Public?

Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation

Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents

To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion

Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy

Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics

In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study

We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Prosody-to-Text: Predicting text from low-pass filtered speech

While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing

We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Backdooring Acoustic Foundation Models for Physically Realizable Triggers

Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated.

Paper