Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR).
Key points
- A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely.
- To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered.
- Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates.
- Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern.
Sources (1)
- [1]Same Outcome, Different Evidence: Intent Recovery in LLM Safety EvaluationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:50 AM
Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR).
A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely.
Extractive summary: sentences quoted from the sources.