The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks.
ProofPaper ↗
Key points
- Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space.
- We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue.
- Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word.
- Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Sources (1)
- [1]The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:17 AM
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks.
Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space.
Extractive summary: sentences quoted from the sources.