ResearchResearch paperSafety & Alignment · Agents & Tool Use · Reinforcement Learning1 source · Oct 8, 2026

One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails

A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text.

Key points

  • Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it.
  • We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost.
  • Escalating the least confident decisions does not help either: a decision an attack has reversed is no less confident than the one it replaced.
  • Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies.

Sources (1)

  • [1]One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:46 PM
    A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text.
    Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
  2. Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
  3. Oct 7, 2026BRANCH: Bypassing Multi-Scanner AI Guardrails
  4. Oct 7, 2026PatchBench: Measuring Collateral Damage in Activation Patching
  5. Oct 7, 2026From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
  6. Oct 7, 2026Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Related