AION
Open-source releaseLarge Language Models · Efficiency & Inference · Training & Scaling1 source · Oct 11, 2026

Converting dense models into Mixture-of-Experts

For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.

Key points

  • If you can turn a dense model into an MoE that only runs part of its MLP per token, you get a model that's cheaper per token for roughly the same knowledge.
  • I have converted two models so far, Qwen/Qwen2.5-0.5B and HuggingFaceTB/SmolLM2-360M (they're purposefully small since my pc can't handle anything else).
  • The conversions can be found here bayliner1980/Qwen2.5-0.5B-MoE-A0.3B and here bayliner1980/SmolLM2-360M-MoE-A0.2B.
  • This does cause a slight increase in MLP compute but I saw it as worth it. bayliner1980/Qwen2.5-0.5B-MoE-A0.3B has one dense at layer 23 and bayliner1980/SmolLM2-360M-MoE-A0.2B has two at layer 3 and layer 31.

Sources (1)

  • [1]Converting dense models into Mixture-of-Experts
    r/LocalLLaMA (top, daily) · Oct 11, 05:45 AM
    For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
    If you can turn a dense model into an MoE that only runs part of its MLP per token, you get a model that's cheaper per token for roughly the same knowledge.

Extractive summary: sentences quoted from the sources.