Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.

Qwen/Qwen-Image-2.1-Turbo
Qwen published the model Qwen-Image-2.1-Turbo on Hugging Face.

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image.
GPT-6 and Intelligent UI for everyone
GPT‐6 is rolling out globally in ChatGPT with Intelligent UI, delivering faster responses with visuals and interactive experiences you can explore and use directly.
mistralai/Voxtral-Mini-4B-Realtime-Arabic
mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.
EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2: an open, lightweight multimodal embedding model
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.

Introducing Mistral Large 4
Today, we’re launching a public preview of Mistral Large 4.

llm-openai-decisions 0.1a0
OpenAI released their new Jev-style Decisions API, as previously announced at last week's DevDay.

Google rolls out improved SynthID AI content detector, now available globally
Every piece of AI content from Google's Gemini models has a hidden SynthID label, which makes it very difficult to pass the content off as authentic.

Odyssey-3 is a new generative world model that you can try for free
Odyssey is making its world model Odyssey-3 available as a public research preview.
nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
nerkyor published the model Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 on Hugging Face.
nvidia/PixelUMM
nvidia published the model PixelUMM on Hugging Face.
Multimodal open d1 decision models for the edge
Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
open-webui/open-webui v0.12.0
Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.
A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding
We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models.
unslothai/unsloth v0.1.904-beta: Train your own Decision model
Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.

Sonnet 5.5 is worth a try
I’ll be at OpenAI DevDay today.
Fully local conversational AI: Whisper + Hermes 8B + Kokoro, zero cloud, running inside a plush toy
A conversational AI plush toy where nothing leaves my network.
Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3
Vision-language models have made it possible to build visual AI agents that understand video at production scale.
Introducing Quine: An AI research system designed for the complexity of biology
Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interactive harness connecting models, scientific tools, literature, and researchers.
microsoft/AesCode-32B
microsoft published the model AesCode-32B on Hugging Face.
[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
Since our founding in 2024, World Labs has built leading AI spatial intelligence capabilities for everything from creative work to design.
huggingface/transformers v5.18.0: Release 5.18.0
Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
ollama/ollama v0.35.1
Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
VideoImage Decision Models for RPA: Forms, Scans and Screenshots
In this video, we return to looking at decision models, but this time for images, with the use case being RPA.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
SuperNav: An Agentic Navigation System for Any Task in Any Scene
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
Reasoning-Informed Visual Editing
To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.
OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
VibeEdit: Image Editing with Canvas Instructions
We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
alesha-pro published the model Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF on Hugging Face.
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
LiquidAI/d1-omni-600M
LiquidAI published the model d1-omni-600M on Hugging Face.
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
V-CoLA: Vision Token Compression with Linear Attention
To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.
SPW-Nav Streams Language-Guided Panoramic Video in Real Time
Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework.
Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.
Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit
Mia-AiLab published the model GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit on Hugging Face.
VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.
On the Necessity of Attention-FFN Split in Vision Transformers
In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs).
Cloudflare/clef-flash
Cloudflare published the model clef-flash on Hugging Face.
Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration.
SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.