Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.

Qwen/Qwen-Image-2.1-Turbo
Qwen published the model Qwen-Image-2.1-Turbo on Hugging Face.

Odyssey-3 is a new generative world model that you can try for free
Odyssey is making its world model Odyssey-3 available as a public research preview.

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image.

llm-openai-decisions 0.1a0
OpenAI released their new Jev-style Decisions API, as previously announced at last week's DevDay.
mistralai/Voxtral-Mini-4B-Realtime-Arabic
mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
GPT-6 and Intelligent UI for everyone
GPT‐6 is rolling out globally in ChatGPT with Intelligent UI, delivering faster responses with visuals and interactive experiences you can explore and use directly.
open-webui/open-webui v0.12.0
Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.

Google rolls out improved SynthID AI content detector, now available globally
Every piece of AI content from Google's Gemini models has a hidden SynthID label, which makes it very difficult to pass the content off as authentic.

Introducing Mistral Large 4
Today, we’re launching a public preview of Mistral Large 4.
EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2: an open, lightweight multimodal embedding model
Fully local conversational AI: Whisper + Hermes 8B + Kokoro, zero cloud, running inside a plush toy
A conversational AI plush toy where nothing leaves my network.
Multimodal open d1 decision models for the edge
Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration.
LiquidAI/d1-omni-600M
LiquidAI published the model d1-omni-600M on Hugging Face.
unslothai/unsloth v0.1.904-beta: Train your own Decision model
Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
Reasoning-Informed Visual Editing
To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
SuperNav: An Agentic Navigation System for Any Task in Any Scene
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
VibeEdit: Image Editing with Canvas Instructions
We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
V-CoLA: Vision Token Compression with Linear Attention
To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework.
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
SPW-Nav Streams Language-Guided Panoramic Video in Real Time
Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.
WOVEN: Weaving Visual World Modeling into Multimodal LLMs
We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs).
From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment
Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content.
DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.
From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections.
Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.
Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning
To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve.
Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation
Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source.
Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.
GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
On the Necessity of Attention-FFN Split in Vision Transformers
In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs).
The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute
We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes.
SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.
No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding
Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm.
4-Tensor Attention Model for Semantic Physical Reality
We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.
DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks.
ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation
We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.