AION
Research paperHardware & Compute · Efficiency & Inference1 source · Oct 7, 2026

Democratizing MoE inference on commodity GPUs with CoMoE

We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.

Key points

  • Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication.
  • Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive.
  • For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects.
  • Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.

Sources (1)

  • [1]Democratizing MoE inference on commodity GPUs with CoMoE
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:24 AM
    We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.
    Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Oct 5, 2026Introducing GLM 5.3 on Amazon Bedrock
  3. Oct 5, 2026Sharing AI progress in mathematics
  4. Sep 28, 2026Holo4: powering generalist computer-use agents
  5. Jun 10, 2026DiffusionGemma: 4x faster text generation
  6. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Related