Visual Memory Attacks Can Persist Through The KV Cache
We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention.
Key points
- Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images.
- Can an adversarial input continue to steer a model even after that input is removed from its context?
- A cache-swap ablation localizes the persistent influence to the KV cache.
- Finally, we show that these attacks can be trained to survive compaction that retains the KV cache of a summary generated by the same model, demonstrating that adversarial behaviour can persist in cached state without continued access to its source.
Sources (1)
- [1]Visual Memory Attacks Can Persist Through The KV CachearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:25 PM
We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention.
Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images.
Extractive summary: sentences quoted from the sources.