AION
Technique

Mixture of Experts

Also known as: MoE, mixture-of-experts

31stories this week
36last 30 days
75all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
    So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
  2. Oct 11, 2026 · Open-source release · 1 source
    Converting dense models into Mixture-of-Experts
    For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
  3. Oct 11, 2026 · Tutorial / explainer · 1 source
    OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!
    I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
  4. Oct 8, 2026 · Research paper · 2 sources
    One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
    In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
  5. Oct 8, 2026 · Research paper · 1 source
    ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
    Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment.
  6. Oct 8, 2026 · Research paper · 2 sources
    MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
    Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
  7. Oct 8, 2026 · Research paper · 1 source
    From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding
    Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm.
  8. Oct 8, 2026 · Research paper · 1 source
    Cost-Aware Mixture-of-Experts Coordination for Model Markets
    This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.
  9. Oct 8, 2026 · Research paper · 1 source
    AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance
    We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling.
  10. Oct 8, 2026 · Research paper · 1 source
    RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
    Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
  11. Oct 8, 2026 · Research paper · 1 source
    Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
    Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.
  12. Oct 8, 2026 · Research paper · 1 source
    WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models
    To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations.
  13. Oct 8, 2026 · Research paper · 1 source
    DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
    DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups.
  14. Oct 8, 2026 · Research paper · 1 source
    When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
    Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.
  15. Oct 7, 2026 · Research paper · 1 source
    When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
    We introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model.
  16. Oct 7, 2026 · Research paper · 1 source
    Democratizing MoE inference on commodity GPUs with CoMoE
    We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.
  17. Oct 7, 2026 · Research paper · 1 source
    Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
    Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.
  18. Oct 7, 2026 · Research paper · 1 source
    Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression
    Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.
  19. Oct 6, 2026 · Research paper · 1 source
    CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
    The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks.
  20. Oct 6, 2026 · Research paper · 1 source
    LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition
    Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index.
  21. Oct 6, 2026 · Research paper · 1 source
    Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
    In this paper, we investigate the mechanics of multi-step arithmetic in compact "Tiny" Transformers ( 10.6M non-embedding parameters, 49.3M total) trained on synthetic data across four basic operations (+, -, , /) unrolled as step-by-step scratchpads.
  22. Oct 6, 2026 · Research paper · 1 source
    A Systematic Study of Small Language Models on Abstract Reasoning Tasks
    Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities.
  23. Oct 6, 2026 · Open-source release · 1 source
    huggingface/transformers v5.19.0: Release v5.19.0
    EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
  24. Oct 6, 2026 · Research paper · 1 source
    Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
    We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes.
  25. Oct 6, 2026 · Research paper · 1 source
    MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling
    We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.
  26. Oct 6, 2026 · Research paper · 1 source
    Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
    Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.
  27. Oct 6, 2026 · Research paper · 1 source
    Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
    Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
  28. Oct 6, 2026 · Research paper · 1 source
    TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
    In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods.
  29. Oct 6, 2026 · Research paper · 1 source
    Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction
    To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM$^2$ES) framework for survival prediction.
  30. Oct 5, 2026 · Product / feature launch · 1 source
    Introducing GLM 5.3 on Amazon Bedrock
    GLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock.

Often appears with