Explore
Everything AION read, in seven sections. Pick one, a topic or a time window.
U-Space: Uncovering When and Why Uncertainty Arises in Language Models
Large language models are informing decisions with ever-higher stakes.
FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment
Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate.
Quoting Ben Affleck
And a tensor, you use a convolutional neural network to identify patterns in that that would reveal what's called edge detection or feature extraction, which is just identifying patterns enough to know like this is where the window ledge is, so we can more easily take the green screen image out and replace it with something.
DSReg: Provably Recovering Individual World Latents without Reconstruction
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity.
MotherTree: Meta-learning on synthetic data improves decision tree training
We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass.
A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding
We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models.
Marformer: A Transformer for Predicting Missing Data Distributions
We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values.
Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields
Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention.
asdex: Automatic Sparse Differentiation in JAX
Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern.
Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study
Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior.
RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
How Do Transformers Learn to Represent Symmetries?
Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning.
Fully Interpretable Minimal Transformers: From Geometry to Algorithm
We present a framework for building and interpreting minimal transformer models.
Leaner Transformers Can Easily Learn to Cluster
Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$).
Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Federated Lightweight Heads, to better capture these irregularities.
Memorization and Malign Generalization in Conditional Diffusion Models with Random Features
Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions.
Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models
Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging.
Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation
PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.
BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment
We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence.
Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.
Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations.
Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).
Stability of Measure-to-Measure Transformers on Sub-Gaussian Data
We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined.
Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.
AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility
We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters.
CARE: Constrained Attention Refinement for Fine-Grained Visual Classification via Teacher-Student Distillation
To address this problem, we propose CARE, a constrained attention refinement framework for interpretable fine-grained recognition via teacher-student distillation.
When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection
We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority.
MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources.
Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction
To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM$^2$ES) framework for survival prediction.
Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have.
Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair
We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data.
Patient, Place, Prior (P$^3$): What Counts as Personalization in Medical World Models?
We introduce Patient, Place, Prior (P$^3$), an audit asking whether a forecast benefits from the patient's longitudinal imaging history (Patient), benefits from patient-matched externally supplied spatial support (Place), and gains predictive value beyond a population-average prediction under matched support and context (Prior).
Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets.
MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge
This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA.
TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks
Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears.
When to Unpair: Regulating Pairing Dependence in Medical Visual In-Context Learning
Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, while the labels collectively indicate the requested task.
Less from More: Reinforcing Sparse Video Reasoning from Dense References
We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference.
Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction.
Few-Shot Learning for Personalised Automated Pain Assessment
In this work, we evaluate Few-Shot Learning, a sub-area of Meta-Learning, as an approach to personalisation in automated pain assessment.
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
Correlational Training of Morphological Neural Networks
In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme.
MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection
In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues.
KGATE : a Knowledge Graph Embedding Training Environment
Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction.
HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data
Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700.
SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection
To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules.
Causal-fate dynamics of unrealized influence
Here we formulate causal-fate dynamics, in which generated influence may be realized, remain latent, or be transformed by subsequent dynamics, and give an exact finite-transport representation when the relevant maps are specified.
Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss
What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning?