From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications.
ProofPaper ↗
Key points
- We establish a precise connection between these two approaches through a probabilistic reformulation.
- Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution.
- Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization.
- Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.
Sources (1)
- [1]From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:40 PM
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications.
We establish a precise connection between these two approaches through a probabilistic reformulation.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- Oct 7, 2026Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
- Oct 6, 2026Secure Speculative Decoding for Large Language Models
- Jul 21, 2026Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber