Research
Papers and datasets worth knowing, ranked by significance and community attention.
Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear.
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.
A self-learning scientific agent for X-ray diffraction
Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM.
Incidental information contaminates patient notes and disrupts clinical reasoning in large language models
We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning.
BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment
We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence.
EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation
We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers.
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.
Leakage-Controlled Multimodal Learning for Diagnosis and Progression Prediction in Alzheimer's Disease Research
Alzheimer's disease prediction involves irregular visits, heterogeneous measurements and incomplete modalities.
Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures
Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices.
InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers
To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory.
TokenBank: Financial Infrastructure for AI Services
We present TokenBank, a financial infrastructure that represents these commitments through structured contracts.
AI-Mediated Self: How HCI Defines and Relates to the Self
This scoping review analyzes 102 papers to examine how the self is defined in the field of human-computer interaction (HCI), how AI-self relationships are conceptualized, and what risks emerge when AI becomes entangled with selfhood.
Revisiting Explainable AI through Model-Independent Concept Dictionaries
To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning.
When Scientific Cognition Is No Longer Scarce
AI could change which parts of science impede progress.
LLM Persuasion Is in the Eye of the Evaluation
Large language models (LLMs) have already been shown to match or exceed human experts in persuasion.
Why Software Engineering Is Indispensable in the Age of Coding Agents
Can AI make Software Engineering (SE) -- the discipline -- obsolete?
When Algorithmic Exploration Becomes Cheap: A Case Study of Agentic Research in EDA
As EDA researchers, we conducted eight deliberate trials of agentic algorithm exploration, selecting several topics outside our areas of depth.
Comprehension Audits to Mitigate Risks from Automated AI Research
We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding.
Healthy skepticism in AI: a data visualization research agenda
Research in data visualization of artificial intelligence (AI) models has historically focused on enhancing trust through visual explanations of AI.
Structured pre-generation elicitation versus single-shot prompting in AI-assisted enterprise decision-making: a randomised online experiment
Generative AI speeds, and mostly improves, professional work, but there is concern that users who delegate both the production and the evaluation of an answer may accept weak output and engage less with the underlying reasoning (cognitive surrender).
The AI Evaluation Ecosystem
We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic actors, building on advances in Generative Agent-Based Modeling (GABM).
Justice After Identity: Large Language Models and the View from Everywhere
Artificial intelligence engaged to calculate algorithmic and agentic fairness introduces a novel possibility.
A Case Study in Assuring AI-Written Software
Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect.
Accelerating the Development of PLGA In Situ Forming Depots Through AI-Driven Multi-Objective Optimization
Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives.
An AI-Assisted Formalization of the Poincaré Conjecture
We present an AI-assisted Lean 4 formalization of the Poincaré conjecture.
MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation
We present MedZERO, a self-evolving framework for open-ended medical reasoning.
Symphony for Text Generation: Benchmarking Clinical Note Generation
We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI.
IEEE 802.11bx - WLAN Intelligent Networking (WIN): Toward an AI-Ready Wi-Fi 9
As a concrete illustration of the AI as traffic paradigm, we present a case study on AI traffic differentiation, where we explore a potential extension of the current Enhanced Distributed Channel Access (EDCA) to support new AI traffic flows.
Evaluating human-AI workflows for field research in viticulture
We assessed the value of two live human-AI interactions in a precision disease control project in California vineyards.