AION
Research paperEvaluation & Benchmarks · Safety & Alignment · Large Language Models1 source · Oct 7, 2026

PatchBench: Measuring Collateral Damage in Activation Patching

To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.

Key points

  • An LLM safety patch can pass a benchmark while still being a poor repair.
  • We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise.
  • For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms.
  • PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.

Sources (1)

  • [1]PatchBench: Measuring Collateral Damage in Activation Patching
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:43 PM
    To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
    An LLM safety patch can pass a benchmark while still being a poor repair.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
  2. Oct 7, 2026Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
  3. Oct 7, 2026Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
  4. Oct 6, 2026Secure Speculative Decoding for Large Language Models
  5. Oct 6, 2026Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
  6. Jul 21, 2026Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Related