ResearchResearch paperLarge Language Models · Multimodal Models1 source · Oct 6, 2026

Selective Transfer of RL Updates for Visual Reasoning

Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training.

Key points

  • We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL).
  • Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update.
  • Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM.
  • Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient.

Sources (1)

  • [1]Selective Transfer of RL Updates for Visual Reasoning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:43 PM
    Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training.
    We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL).

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  2. Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
  3. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  4. Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
  5. Oct 1, 2026nvidia/PixelUMM
  6. Sep 28, 2026Holo4: powering generalist computer-use agents

Related