Research
Papers and datasets worth knowing, ranked by significance and community attention.
OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
Marformer: A Transformer for Predicting Missing Data Distributions
We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values.
Task-Sufficient Contraction: Source Selection for Machine Information Interfaces
This paper studies when one such reduction preserves the complete downstream problem family, a property termed Task-Sufficient Contraction.
Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic.
Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions
We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews.
RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees
Retrieval-augmented generation (RAG) grounds language models in external corpora.
RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation.
EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation
Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable.
From Chunks to Functional Evidence: Function-Aware Retrieval for EDA Documentation QA
Retrieval-Augmented Generation (RAG) is widely used to ground answers in documents.
SIFT: Search Intent-to-Filter Transformer for Multi-Task Personalized Filter Ranking at Airbnb
We present SIFT (Search Intent-to-Filter Transformer), a ranking model built on transformers that learns guest preferences directly from raw behavioral sequences.
Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear.
EC-RAG: Event Chain Retrieval-Augmented Generation for Long Video Understanding
In this paper, we propose Event Chain Retrieval-Augmented Generation (EC-RAG), a training-free framework that organizes video content into an explicit event chain before question answering.
RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems
We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits.
Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents
We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution.
Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek
We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes.
Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation
Varying the forced reasoning language of an agentic RAG system on this testbed, we find that aligning the reasoning language with the language of the query and the retrieved documents helps.
Compile the Table: Query-Calibrated Operator Compression for Tabular In-Context Learning
We propose QCOC (Query-Calibrated Operator Compression), which exploits the exchangeability and repeated use of in-context examples by compiling their full KV cache once into compact memory shared across subsequent queries.
Learning to Retrieve via Reinforcement Learning in Embedding Space
To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards.
Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval
Generative retrieval trains a language model to generate the identifier of a relevant document.
Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have.
RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models
Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies.
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
KGATE : a Knowledge Graph Embedding Training Environment
Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction.
UNREAL: Unifying Retrieval and Long-Context with a Single Model
We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference.
NativeScope: Relation-Localized Retrieval over Native Topology with a Correct Anchor
We propose NativeScope, a scope-then-rank method for queries with a known anchor and relation.
Training with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition
Generative recommenders return a limited candidate set and may omit observed targets before reranking.
Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers.
From Retrieval to Reconstruction: Constructing Evolvable Cognitive Memory for Long-Term Dialogue
Large Language Models (LLMs) serving as long-term dialogue agents require memory systems that support reliable reasoning over extended interactions.
Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis
We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist.
Thinking in Depth: Retrospective Inference for Tabular Foundation Models
We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network.
ExperienceIndex: Artifact-Grounded Memory
We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces.
Beyond Resolution: Object-to-Image Ratio Mismatch in Instance Retrieval
We show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images.
Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora
Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval.
Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources.
Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval.
Embedding-Bias in Conditional Independence Testing
To test conditional independence of $X$ and $Y$ given a text or an image $Z$, one conditions on an embedding $ψ(Z)$ in place of $Z$.
Beyond Sequences: Distilling Structured Decision Memory for LLM Recommendation
To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence.
Efficient Multi-Granularity Knowledge Transfer for Radiology Report Generation
Radiology report generation can automatically generate clinical descriptions from X-ray images, thereby significantly improving the efficiency of radiologists.
SteerCast: Retrieval-Based Latent Steering for Decoder-Only Time Series Forecasting
We propose SteerCast, a retrieval-based latent steering method that improves decoder-only forecaster at inference time, without updating its parameters.
Multimodal Graph Retrieval-Augmented Sequential Recommendation via Collaborative Filtering Paths
To address these challenges, we propose MGRASRec, a multimodal graph retrieval-augmented framework for sequential recommendation.
Finding the Right Balance: Relevance and Diversity in LLM Retrieval
Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages.
Contrastive Learning for Aspect Representation towards Explainable Recommendation
In this work, we propose a novel recommendation model, CLARER (Contrastive Learning for Aspect Representation towards Explainable Recommendation) that integrates aspect features learned from textual reviews with rating information to improve the accuracy and explainability of recommendations.
From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis
We define a customer context graph as a unified model of customer and business context.
TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning
We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents.
Lossy Compressive Text Autoencoders
Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning.
Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models
Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt.
Trustworthy Domain-Specific AI for Structured Knowledge Retrieval and Reasoning
This dissertation presents a scalable architecture for transforming unstructured, domain-specific text into structured knowledge for retrieval and reasoning.