Democratizing MoE inference on commodity GPUs with CoMoE
We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.
Key points
- Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication.
- Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive.
- For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects.
- Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.
Sources (1)
- [1]Democratizing MoE inference on commodity GPUs with CoMoEarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:24 AM
We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.
Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Oct 5, 2026Introducing GLM 5.3 on Amazon Bedrock
- Oct 5, 2026Sharing AI progress in mathematics
- Sep 28, 2026Holo4: powering generalist computer-use agents
- Jun 10, 2026DiffusionGemma: 4x faster text generation
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model