ResearchResearch paperSafety & Alignment1 source · Oct 7, 2026

From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents

When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications.

Key points

  • We establish a precise connection between these two approaches through a probabilistic reformulation.
  • Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution.
  • Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization.
  • Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
  2. Oct 7, 2026Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
  3. Oct 6, 2026Secure Speculative Decoding for Large Language Models
  4. Jul 21, 2026Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Related