Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

HF: Qwen4 sources9h ago

Qwen/Qwen-Image-2.1-Turbo

Qwen published the model Qwen-Image-2.1-Turbo on Hugging Face.

Weights
Cloudflare Blog: AI2 sources2d ago

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Today, we’re releasing Clef-omni, which takes in audio and video input alongside text and image.

Weights
OpenAI News2 sources4d ago

GPT-6 and Intelligent UI for everyone

GPT‐6 is rolling out globally in ChatGPT with Intelligent UI, delivering faster responses with visuals and interactive experiences you can explore and use directly.

HF: Mistral AI3d ago

mistralai/Voxtral-Mini-4B-Realtime-Arabic

mistralai published the model Voxtral-Mini-4B-Realtime-Arabic on Hugging Face.

Weights
Google DeepMind Blog2 sources5d ago

EmbeddingGemma 2: an open, lightweight multimodal embedding model

EmbeddingGemma 2: an open, lightweight multimodal embedding model

Paper
Hugging Face Daily Papers2 sources3d ago

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.

▲ 52 upvotesPaper
Mistral AI News3 sources4d ago

Introducing Mistral Large 4

Today, we’re launching a public preview of Mistral Large 4.

2 outlets
Simon Willison's Weblog2 sources2d ago

llm-openai-decisions 0.1a0

OpenAI released their new Jev-style Decisions API, as previously announced at last week's DevDay.

2 outlets
Ars Technica: AI2 sources3d ago

Google rolls out improved SynthID AI content detector, now available globally

Every piece of AI content from Google's Gemini models has a hidden SynthID label, which makes it very difficult to pass the content off as authentic.

2 outlets
The Decoder2 sources13h ago

Odyssey-3 is a new generative world model that you can try for free

Odyssey is making its world model Odyssey-3 available as a public research preview.

2 outlets
Hugging Face trending models7d ago

nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2

nerkyor published the model Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 on Hugging Face.

Weights
HF: NVIDIA9d ago

nvidia/PixelUMM

nvidia published the model PixelUMM on Hugging Face.

Weights
Hugging Face Blog4d ago

Multimodal open d1 decision models for the edge

Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).

Vendor claim only
GitHub: open-webui/open-webui1d ago

open-webui/open-webui v0.12.0

Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.

Code
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding

We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models.

Paper
GitHub: unslothai/unsloth4d ago

unslothai/unsloth v0.1.904-beta: Train your own Decision model

Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.

Code
Ben's Bites12d ago

Sonnet 5.5 is worth a try

I’ll be at OpenAI DevDay today.

r/LocalLLaMA (top, daily)1d ago

Fully local conversational AI: Whisper + Hermes 8B + Kokoro, zero cloud, running inside a plush toy

A conversational AI plush toy where nothing leaves my network.

NVIDIA Technical Blog12d ago

Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3

Vision-language models have made it possible to build visual AI agents that understand video at production scale.

Vendor claim only
Microsoft Research Blog12d ago

Introducing Quine: An AI research system designed for the complexity of biology

Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interactive harness connecting models, scientific tools, literature, and researchers.

Vendor claim only
HF: Microsoft12d ago

microsoft/AesCode-32B

microsoft published the model AesCode-32B on Hugging Face.

Weights
Latent Space12d ago

[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more

Since our founding in 2024, World Labs has built leading AI spatial intelligence capabilities for everything from creative work to design.

GitHub: huggingface/transformers11d ago

huggingface/transformers v5.18.0: Release 5.18.0

Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.

Code
GitHub: ollama/ollama12d ago

ollama/ollama v0.35.1

Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.

Code
Video
Sam Witteveen (YouTube)9d ago

Image Decision Models for RPA: Forms, Scans and Screenshots

In this video, we return to looking at decision models, but this time for images, with the use case being RPA.

Paper
Hugging Face Daily Papers2 sources3d ago

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

SuperNav: An Agentic Navigation System for Any Task in Any Scene

General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

Reasoning-Informed Visual Editing

To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video

We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

VibeEdit: Image Editing with Canvas Instructions

We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.

Paper
Hugging Face trending models7d ago

alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF

alesha-pro published the model Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF on Hugging Face.

Weights
Paper
Hugging Face Daily Papers2 sources4d ago

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.

Paper
Hugging Face trending models6d ago

LiquidAI/d1-omni-600M

LiquidAI published the model d1-omni-600M on Hugging Face.

Weights
Paper
Hugging Face Daily Papers2 sources3d ago

Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models

Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

V-CoLA: Vision Token Compression with Linear Attention

To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

SPW-Nav Streams Language-Guided Panoramic Video in Real Time

Researchers introduced SPW-Nav, a panoramic world model that interprets language movement instructions to stream one minute of real-time 2K 360-degree video from a single panorama.

Paper
Paper
Hugging Face Daily Papers2 sources3d ago

OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework.

Paper
Paper
Hugging Face Daily Papers2 sources4d ago

Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.

Paper
Hugging Face trending models6d ago

Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit

Mia-AiLab published the model GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit on Hugging Face.

Weights
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

On the Necessity of Attention-FFN Split in Vision Transformers

In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs).

Paper
Hugging Face trending models11d ago

Cloudflare/clef-flash

Cloudflare published the model clef-flash on Hugging Face.

Weights
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings

To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.

Paper