AION
Research paperSafety & Alignment · Large Language Models1 source · Oct 6, 2026

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.

Key points

  • As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering.
  • Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety.
  • We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not.
  • Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing.

Sources (1)

Extractive summary: sentences quoted from the sources.