PatchBench: Measuring Collateral Damage in Activation Patching
To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
Key points
- An LLM safety patch can pass a benchmark while still being a poor repair.
- We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise.
- For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms.
- PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.
Sources (1)
- [1]PatchBench: Measuring Collateral Damage in Activation PatchingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:43 PM
To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
An LLM safety patch can pass a benchmark while still being a poor repair.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
- Oct 7, 2026Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- Oct 7, 2026Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
- Oct 6, 2026Secure Speculative Decoding for Large Language Models
- Oct 6, 2026Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
- Jul 21, 2026Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber