AION
Research paperLarge Language Models · Safety & Alignment1 source · Oct 8, 2026

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR).

Key points

  • A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely.
  • To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered.
  • Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates.
  • Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern.

Sources (1)

  • [1]Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:50 AM
    Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR).
    A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely.

Extractive summary: sentences quoted from the sources.

Related