Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
Key points
- As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering.
- Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety.
- We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not.
- Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing.
Sources (1)
- [1]Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:09 AM
Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering.
Extractive summary: sentences quoted from the sources.