Prompt injection
Also known as: prompt injections
8stories this week
8last 30 days
10all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceOne Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent GuardrailsA typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text.
- Oct 8, 2026 · Research paper · 1 sourceLTBD: Learnable Trust-Boundary Delimiters for Prompt Injection DefenseTo address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged.
- Oct 7, 2026 · Research paper · 1 sourceBRANCH: Bypassing Multi-Scanner AI GuardrailsWe propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems.
- Oct 7, 2026 · Research paper · 1 sourcePackage Hallucination Attacks on Coding Agents through Prompt Injection in Rule FilesTo bridge this gap, we introduce the package hallucination attack, where an attacker injects malicious prompts into benign rule files to induce coding agents to replace legitimate dependencies with attacker-controlled packages.
- Oct 6, 2026 · Research paper · 1 sourceAdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World ModelWe introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model.
- Oct 6, 2026 · Research paper · 1 sourceSecure Speculative Decoding for Large Language ModelsSpeculative decoding accelerates inference for a large language model (LLM), referred to as the target model, by first using a smaller model, referred to as the draft model, to generate candidate tokens and then verifying them with the target model for acceptance or rejection.
- Oct 6, 2026 · Research paper · 1 sourceRAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG SystemsWe introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits.
- Oct 6, 2026 · Research paper · 1 sourceSurviving the Router: Optimizing Skill Injections for Retrieval and ExecutionTo address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks.
- Jun 24, 2026 · Product / feature launch · 1 sourceIntroducing computer use in Gemini 3.5 FlashIntroducing computer use in Gemini 3.5 Flash
- Jun 16, 2026 · Opinion / analysis · 1 sourceSecuring the future of AI agentsHow we’re securing internal systems against increasingly capable and imperfectly aligned AI