MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling
We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.
ProofPaper ↗
Key points
- Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation.
- Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts.
- Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine.
- These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.
Sources (1)
- [1]MASKerade: Token-Routed Mask Experts for Dense-to-MoE UpcyclingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:04 AM
We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.
Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 15, 2026vllm-project/vllm v0.23.0
- Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model