Research
Papers and datasets worth knowing, ranked by significance and community attention.
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution.
From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope.
Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.
Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user.
How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.
TestPrism: Rethinking Test Evaluation Beyond a Single Reference
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks.
Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Here, we develop an ecological theory of AI-agent populations based on a population growth equation in which fitness (growth rate) depends on cybersecurity capability.
Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic.
Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs
Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated.
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.
Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents
We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance.
On the estimation and validity of AI time horizons---a statistical look at the METR plot
On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time.
A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration
Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots.
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.
LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged.
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment.
EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation
We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers.
Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation
In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes.
Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper
We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text.
Traceable World State: A Provenance-Aware State Representation and Deterministic Replay Framework for Robotic Systems
We present Traceable World State (TWS), a middleware-neutral semantic representation and reference runtime for provenance-aware robot world state.
Distillation for Incrimination and Distillation for Capabilities
Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative.
Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents
Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks and keystrokes.
SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing.
Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study
We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence.
World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models
Robot world models support policy evaluation, planning, and synthetic data generation, but these applications require predictions that distinguish successful actions from failures.
What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents
To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph.
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants.
PatchBench: Measuring Collateral Damage in Activation Patching
To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
To bridge this gap, we introduce the package hallucination attack, where an attacker injects malicious prompts into benign rule files to induce coding agents to replace legitimate dependencies with attacker-controlled packages.
RH-Detect: A Unified Benchmark for Reward Hacking Detection
We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology.
DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists
We introduce DrugTargetWorld, a framework that procedurally generates simulated biobanks, or "worlds," with known but concealed causal structure.
From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications.
FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training.
OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step.
Large-scale Repository Engineering via Agent-Native Reusable Code Primitives
We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests.
When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents
We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures.
HE-OFT: Privacy-Preserving One-Shot Federated Fine-Tuning under Homomorphic Encryption
We present HE-OFT, the first cryptographically secure one-shot federated fine-tuning protocol in which no party receives the trained model.
Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models
In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations.
AgentEvolver: System-Wide Self-Evolution Through Task Execution
We present AgentEvolver, a system for developing capabilities during task execution while keeping the foundation model fixed.
Chronos Enables Code Agents to Reason over Software Evolution
We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents.
Does On-Policy Distillation for Safety Pose Backdoor Risks?
On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models.
How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern.