Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions
To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection.
ProofPaper ↗
Key points
- Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks.
- Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking.
- We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps.
- Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead.
Sources (1)
- [1]Purifying Backdoored Large Vision-Language Models by Removing Hijacked DirectionsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:25 PM
To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection.
Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching
- Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 28, 2026openai/openai-python v3.20.0
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0