When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
We introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model.
Key points
- Mixture-of-Experts (MoE) language models produce routing information during inference that may be logged or exposed for monitoring, debugging, load analysis, and safety auditing.
- Unlike ordinary model outputs, this telemetry reveals a view of the model's internal computation, raising a privacy question: can it reveal whether an example was used to fine-tune the deployed model?
- Across three MoE architectures and three data domains, router telemetry consistently improves membership inference over a strong output-signal ensemble, increasing TPR at 1% FPR by 2.7--9.4 percentage points across all nine settings.
- Mechanistic analysis further shows that the leakage does not require router-specific memorization: fine-tuning introduces membership information into hidden representations, while the router exposes a projection of this signal even when its parameters are frozen.
Sources (1)
- [1]When Routing Reveals Membership: Privacy Leakage from MoE Router TelemetryarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:26 AM
We introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model.
Mixture-of-Experts (MoE) language models produce routing information during inference that may be logged or exposed for monitoring, debugging, load analysis, and safety auditing.
Extractive summary: sentences quoted from the sources.