AION
Research paperReinforcement Learning · Robotics & Embodied AI1 source · Oct 7, 2026

RFPO: Rectified Flow Policy Optimization for Embodied Control

Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost.

Key points

  • To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution.
  • Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration.
  • A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step.
  • Real-robot experiments further validate stable one-step locomotion.

Sources (1)

  • [1]RFPO: Rectified Flow Policy Optimization for Embodied Control
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:24 PM
    Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost.
    To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution.

Extractive summary: sentences quoted from the sources.