ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents.
ProofPaper ↗
Key points
- As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important.
- Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors.
- Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence.
- We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation.
Sources (1)
- [1]ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic CodingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:46 PM
To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents.
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important.
Extractive summary: sentences quoted from the sources.