Mixture of Experts
Also known as: MoE, mixture-of-experts
31stories this week
36last 30 days
75all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceQwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysisSo far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
- Oct 11, 2026 · Open-source release · 1 sourceConverting dense models into Mixture-of-ExpertsFor the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceOMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
- Oct 8, 2026 · Research paper · 2 sourcesOne Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsIn this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
- Oct 8, 2026 · Research paper · 1 sourceReSI: Recursive Safety Improvement toward Resistant and Resilient AIRecursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment.
- Oct 8, 2026 · Research paper · 2 sourcesMiMo-V2.6: Scaling Reinforcement Learning Towards Self-ImprovementReinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
- Oct 8, 2026 · Research paper · 1 sourceFrom Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion UnderstandingInspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm.
- Oct 8, 2026 · Research paper · 1 sourceCost-Aware Mixture-of-Experts Coordination for Model MarketsThis paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism.
- Oct 8, 2026 · Research paper · 1 sourceAuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert GuidanceWe present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling.
- Oct 8, 2026 · Research paper · 1 sourceRouterInterp: Understanding Superposed Specialisation in Mixture of Experts RoutingLeveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
- Oct 8, 2026 · Research paper · 1 sourceSmoothing the Top-k Exposure Boundary for Sparse Mixture-of-ExpertsSparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.
- Oct 8, 2026 · Research paper · 1 sourceWARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action ModelsTo address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations.
- Oct 8, 2026 · Research paper · 1 sourceDivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert CompositionDivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups.
- Oct 8, 2026 · Research paper · 1 sourceWhen Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM QuantizationMotivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.
- Oct 7, 2026 · Research paper · 1 sourceWhen Routing Reveals Membership: Privacy Leakage from MoE Router TelemetryWe introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model.
- Oct 7, 2026 · Research paper · 1 sourceDemocratizing MoE inference on commodity GPUs with CoMoEWe present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.
- Oct 7, 2026 · Research paper · 1 sourceExpert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token ShufflingMixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.
- Oct 7, 2026 · Research paper · 1 sourceShared Low-rank Basis Factorization for Data-free Mixture-of-Experts CompressionMixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.
- Oct 6, 2026 · Research paper · 1 sourceCurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted SearchThe best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks.
- Oct 6, 2026 · Research paper · 1 sourceLayerRoPE: Dynamic Depth-wise Magnitude & Angular SuperpositionAcross 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index.
- Oct 6, 2026 · Research paper · 1 sourceAlgorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny TransformersIn this paper, we investigate the mechanics of multi-step arithmetic in compact "Tiny" Transformers ( 10.6M non-embedding parameters, 49.3M total) trained on synthetic data across four basic operations (+, -, , /) unrolled as step-by-step scratchpads.
- Oct 6, 2026 · Research paper · 1 sourceA Systematic Study of Small Language Models on Abstract Reasoning TasksEndpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities.
- Oct 6, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.19.0: Release v5.19.0EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
- Oct 6, 2026 · Research paper · 1 sourceBeyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program RepairWe propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes.
- Oct 6, 2026 · Research paper · 1 sourceMASKerade: Token-Routed Mask Experts for Dense-to-MoE UpcyclingWe introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.
- Oct 6, 2026 · Research paper · 1 sourceAdaptive Mean Estimation by In-Context Learning: A Gradient-Flow AnalysisPrior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.
- Oct 6, 2026 · Research paper · 1 sourceReading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language ModelsVision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs.
- Oct 6, 2026 · Research paper · 1 sourceTRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language ModelsIn this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods.
- Oct 6, 2026 · Research paper · 1 sourceAnchor-driven Multi-modal Multi-scale Expert Selection for Survival PredictionTo address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM$^2$ES) framework for survival prediction.
- Oct 5, 2026 · Product / feature launch · 1 sourceIntroducing GLM 5.3 on Amazon BedrockGLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock.