Optimal Transport Q-Learning for Flow Policy Steering and Acceleration
For roboticists using flow-based policies, OTQL provides a sample-efficient method to improve policy performance and accelerate inference without costly distillation.
OTQL fine-tunes suboptimal flow-based policies using RL post-training with only 50-60 episodes of interaction, increasing success rates from 36% to 86% for single-task policies and from 38% to 76% for a pre-trained VLA, while reducing inference steps by 70%.
Diffusion and flow policies have recently demonstrated remarkable performance in robotic applications by accurately capturing multimodal robot trajectory distributions, especially in the context of vision language action (VLA) models. However, high quality policy performance also requires fast inference and high quality demonstrations, which are often hard to get. Lack of these leads to suboptimal policy behaviors and failure under distribution shifts. In this work we address the problem of fine-tuning and accelerating suboptimal flow-based policies using the robot's experience through RL post-training. We introduce Optimal Transport Q-Learning (OTQL), a new method for finetuning flow policies using advantage weighted conditional optimal transport flow matching. OTQL can finetune and accelerate flows with an interaction budget of 50-60 episodes while avoiding computationally expensive distillation in simulation and real-world robot tasks. Our results show that OTQL post-trains flow policies using the robot's own experience, increasing average success percentage of single-task policies from 36% to 86% and of a pre-trained VLA from 38% to 76% while reducing the number of inference steps per action generation by 70%.