BRANCH: Bypassing Multi-Scanner AI Guardrails
We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems.
Key points
- AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks.
- In response, guardrail systems formed by multiple scanners have emerged that collaboratively detect different types of malicious instructions, whereby shared latent representations across classification boundaries render established bypassing techniques ineffective.
- Our method leverages a branching tree search approach that dynamically applies adversarial perturbation against individual scanners, with subsequent perturbation optimization and technique selection based on overall improvement across all guardrail system scanners, effectively decoupling bypass evaluation from attack signal optimization.
- Our findings demonstrate that BRANCH achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques, while preserving semantic meaning within the bypass.
Sources (1)
- [1]BRANCH: Bypassing Multi-Scanner AI GuardrailsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:10 PM
We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems.
AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026PatchBench: Measuring Collateral Damage in Activation Patching
- Oct 7, 2026From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
- Oct 7, 2026Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
- Oct 6, 2026AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
- Oct 6, 2026Secure Speculative Decoding for Large Language Models
- Oct 6, 2026RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems