AION
Concept

Jailbreak

Also known as: jailbreaking, jailbreaks

7stories this week
7last 30 days
8all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
    A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text.
  2. Oct 7, 2026 · Research paper · 1 source
    BRANCH: Bypassing Multi-Scanner AI Guardrails
    We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems.
  3. Oct 7, 2026 · Research paper · 1 source
    PatchBench: Measuring Collateral Damage in Activation Patching
    To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
  4. Oct 7, 2026 · Research paper · 1 source
    From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
    When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications.
  5. Oct 7, 2026 · Research paper · 1 source
    Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
    Looped Language Models (LoopLMs) provide a parameter efficient approach to scaling model capabilities through repeated use of shared parameters across recurrent steps.
  6. Oct 7, 2026 · Research paper · 1 source
    Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
    Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored.
  7. Oct 6, 2026 · Research paper · 1 source
    Secure Speculative Decoding for Large Language Models
    Speculative decoding accelerates inference for a large language model (LLM), referred to as the target model, by first using a smaller model, referred to as the draft model, to generate candidate tokens and then verifying them with the target model for acceptance or rejection.
  8. Jul 21, 2026 · Product / feature launch · 1 source
    Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
    Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Often appears with