ResearchResearch paperTraining & Scaling · Efficiency & Inference · Large Language Models1 source · Oct 6, 2026

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.

Key points

  • Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation.
  • Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts.
  • Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine.
  • These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.

Sources (1)

  • [1]MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:04 AM
    We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.
    Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
  2. Aug 10, 2026vllm-project/vllm v0.27.0
  3. Jun 29, 2026vllm-project/vllm v0.24.0
  4. Jun 15, 2026vllm-project/vllm v0.23.0
  5. Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0
  6. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Related