AION
Research paperLarge Language Models · Multimodal Models · Reinforcement Learning1 source · Oct 7, 2026

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).

Key points

  • Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward.
  • We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases.
  • We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation.
  • Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B.

Sources (1)

  • [1]SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:42 PM
    Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR).
    Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
  2. Oct 7, 2026RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
  3. Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
  4. Oct 7, 2026Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
  5. Oct 6, 2026FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
  6. Sep 30, 2026huggingface/transformers v5.18.0: Release 5.18.0

Related