ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks.

Key points

  • Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space.
  • We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue.
  • Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word.
  • Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related