One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text.
ProofPaper ↗
Key points
- Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it.
- We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost.
- Escalating the least confident decisions does not help either: a decision an attack has reversed is no less confident than the one it replaced.
- Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies.
Sources (1)
- [1]One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent GuardrailsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:46 PM
A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text.
Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
- Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
- Oct 7, 2026BRANCH: Bypassing Multi-Scanner AI Guardrails
- Oct 7, 2026PatchBench: Measuring Collateral Damage in Activation Patching
- Oct 7, 2026From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
- Oct 7, 2026Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files