DecepEval: A Benchmark for Evaluating Deception in LLM Agents
As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment.
ProofPaper ↗
Key points
- To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios.
- Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict.
- DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors.
- Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates.
Sources (1)
- [1]DecepEval: A Benchmark for Evaluating Deception in LLM AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:33 AM
As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment.
To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios.
Extractive summary: sentences quoted from the sources.