How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.
Key points
- We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese.
- It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes.
- Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away.
- We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response.
Sources (1)
- [1]How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and DefensearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:40 PM
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper.
We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation
- Oct 6, 2026Image Bitstream Fine-grained Understanding for Privacy-Friendly AIoT
- Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
- Oct 2, 2026sgl-project/sglang v0.5.21