MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
Key points
- This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute.
- Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up.
- We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency.
- We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
Sources (2)
- [1]MiMo-V2.6: Scaling Reinforcement Learning Towards Self-ImprovementHugging Face Daily Papers · Oct 8, 12:00 AM
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement.
This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute.
- [2]MiMo-V2.6: Scaling Reinforcement Learning Towards Self-ImprovementarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:41 PM · same content
Extractive summary: sentences quoted from the sources.