Q-Learning with Scalar Adjoint Matching
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
Key points
- We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal.
- Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products.
- Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions.
- To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot.
Sources (2)
- [1]Q-Learning with Scalar Adjoint MatchingHugging Face Daily Papers · Oct 7, 12:00 AM
Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal.
- [2]Q-Learning with Scalar Adjoint MatchingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:13 PM · same content
Extractive summary: sentences quoted from the sources.