Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Paper
Hugging Face Daily Papers2 sources5d ago

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Large language models are informing decisions with ever-higher stakes.

▲ 38 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment

Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate.

Paper
Simon Willison's Weblog3d ago

Quoting Ben Affleck

And a tensor, you use a convolutional neural network to identify patterns in that that would reveal what's called edge detection or feature extraction, which is just identifying patterns enough to know like this is where the window ledge is, so we can more easily take the green screen image out and replace it with something.

Paper
Hugging Face Daily Papers2 sources4d ago

DSReg: Provably Recovering Individual World Latents without Reconstruction

Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity.

▲ 3 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MotherTree: Meta-learning on synthetic data improves decision tree training

We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding

We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Marformer: A Transformer for Predicting Missing Data Distributions

We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields

Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

asdex: Automatic Sparse Differentiation in JAX

Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study

Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

How Do Transformers Learn to Represent Symmetries?

Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Fully Interpretable Minimal Transformers: From Geometry to Algorithm

We present a framework for building and interpreting minimal transformer models.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Leaner Transformers Can Easily Learn to Cluster

Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Temporal transformer CAN encoder with federated lightweight heads for anomaly detection

Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Memorization and Malign Generalization in Conditional Diffusion Models with Random Features

Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models

Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation

To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation

PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment

We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence.

Paper
Paper
Hugging Face Daily Papers3d ago

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers

Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?

A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility

We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

CARE: Constrained Attention Refinement for Fine-Grained Visual Classification via Teacher-Student Distillation

To address this problem, we propose CARE, a constrained attention refinement framework for interpretable fine-grained recognition via teacher-student distillation.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection

We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification

We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction

To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM$^2$ES) framework for survival prediction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair

We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Patient, Place, Prior (P$^3$): What Counts as Personalization in Medical World Models?

We introduce Patient, Place, Prior (P$^3$), an audit asking whether a forecast benefits from the patient's longitudinal imaging history (Patient), benefits from patient-matched externally supplied spatial support (Place), and gains predictive value beyond a population-average prediction under matched support and context (Prior).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge

This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks

Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

When to Unpair: Regulating Pairing Dependence in Medical Visual In-Context Learning

Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, while the labels collectively indicate the requested task.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Less from More: Reinforcing Sparse Video Reasoning from Dense References

We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Few-Shot Learning for Personalised Automated Pain Assessment

In this work, we evaluate Few-Shot Learning, a sub-area of Meta-Learning, as an approach to personalisation in automated pain assessment.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must

This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Correlational Training of Morphological Neural Networks

In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection

In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

KGATE : a Knowledge Graph Embedding Training Environment

Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection

To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Causal-fate dynamics of unrealized influence

Here we formulate causal-fate dynamics, in which generated influence may be realized, remain latent, or be transformed by subsequent dynamics, and give an exact finite-transport representation when the relevant maps are specified.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning?

Paper