Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
Hugging Face Daily Papers2 sources5d ago

Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear.

▲ 8 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.

Paper
Paper
Hugging Face Daily Papers2 sources5d ago

Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.

▲ 6 upvotesPaper
Paper
Hugging Face Daily Papers2 sources5d ago

A self-learning scientific agent for X-ray diffraction

Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM.

▲ 17 upvotesPaper
Paper
Hugging Face Daily Papers2 sources5d ago

Incidental information contaminates patient notes and disrupts clinical reasoning in large language models

We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning.

▲ 6 upvotesPaper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment

We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation

We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Leakage-Controlled Multimodal Learning for Diagnosis and Progression Prediction in Alzheimer's Disease Research

Alzheimer's disease prediction involves irregular visits, heterogeneous measurements and incomplete modalities.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers

To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

TokenBank: Financial Infrastructure for AI Services

We present TokenBank, a financial infrastructure that represents these commitments through structured contracts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

AI-Mediated Self: How HCI Defines and Relates to the Self

This scoping review analyzes 102 papers to examine how the self is defined in the field of human-computer interaction (HCI), how AI-self relationships are conceptualized, and what risks emerge when AI becomes entangled with selfhood.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Revisiting Explainable AI through Model-Independent Concept Dictionaries

To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

When Scientific Cognition Is No Longer Scarce

AI could change which parts of science impede progress.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

LLM Persuasion Is in the Eye of the Evaluation

Large language models (LLMs) have already been shown to match or exceed human experts in persuasion.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Why Software Engineering Is Indispensable in the Age of Coding Agents

Can AI make Software Engineering (SE) -- the discipline -- obsolete?

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

When Algorithmic Exploration Becomes Cheap: A Case Study of Agentic Research in EDA

As EDA researchers, we conducted eight deliberate trials of agentic algorithm exploration, selecting several topics outside our areas of depth.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Comprehension Audits to Mitigate Risks from Automated AI Research

We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Healthy skepticism in AI: a data visualization research agenda

Research in data visualization of artificial intelligence (AI) models has historically focused on enhancing trust through visual explanations of AI.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Structured pre-generation elicitation versus single-shot prompting in AI-assisted enterprise decision-making: a randomised online experiment

Generative AI speeds, and mostly improves, professional work, but there is concern that users who delegate both the production and the evaluation of an answer may accept weak output and engage less with the underlying reasoning (cognitive surrender).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

The AI Evaluation Ecosystem

We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic actors, building on advances in Generative Agent-Based Modeling (GABM).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Justice After Identity: Large Language Models and the View from Everywhere

Artificial intelligence engaged to calculate algorithmic and agentic fairness introduces a novel possibility.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

A Case Study in Assuring AI-Written Software

Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Accelerating the Development of PLGA In Situ Forming Depots Through AI-Driven Multi-Objective Optimization

Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

An AI-Assisted Formalization of the Poincaré Conjecture

We present an AI-assisted Lean 4 formalization of the Poincaré conjecture.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

We present MedZERO, a self-evolving framework for open-ended medical reasoning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Symphony for Text Generation: Benchmarking Clinical Note Generation

We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

IEEE 802.11bx - WLAN Intelligent Networking (WIN): Toward an AI-Ready Wi-Fi 9

As a concrete illustration of the AI as traffic paradigm, we present a case study on AI traffic differentiation, where we explore a potential extension of the current Enhanced Distributed Channel Access (EDCA) to support new AI traffic flows.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Evaluating human-AI workflows for field research in viticulture

We assessed the value of two live human-AI interactions in a precision disease control project in California vineyards.

Paper