ROJun 22

Flowing With Purpose: Latent Action Guided Flow Matching Policies For Robotic Manipulation

arXiv:2606.234205.1
Predicted impact top 74% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For robotic manipulation, LAFM addresses a structural limitation in flow matching policies, yielding substantial performance gains and surpassing larger pre-trained models.

Flow matching policies for robotic manipulation suffer from using a fixed isotropic source distribution mismatched to fragmented action spaces. The proposed Latent Action Guided Flow Matching (LAFM) replaces this with an adaptive library of learned priors, improving task success by 23.4% in real-world deployments and 10.4% on LIBERO-90, achieving SOTA with smaller models.

Flow matching has recently become a new standard for behavior cloning in robotic manipulation. However, state-of-the-art flow matching policies suffer from a systematic structural mismatch: they rely on a globally fixed isotropic source distribution despite the strongly fragmented and heteroscedastic structure of robotic action spaces. This agnostic initialization forces the model to learn highly entangled vector fields, bottlenecking training efficiency and limiting overall policy performance. To address this limitation, we introduce Latent Action Guided Flow Matching (LAFM), a novel framework that replaces the monolithic Gaussian with an adaptive library of learned prior distributions. By grounding these distributions using a latent action model, LAFM maps current observations to discrete motion primitives, selecting a specialized base distribution that provides an informed, structurally aligned initialization for the denoising process. This dynamic adaptivity naturally accommodates heteroscedasticity in human demonstrations and makes transport trajectories shorter and less entangled. Empirically, LAFM substantially outperforms standard flow matching formulations, increasing task success rates by 23.4% in real-world robotic deployments and by 10.4% on the LIBERO-90 benchmark. Furthermore, we demonstrate that LAFM achieves state-of-the-art results, surpassing massively pre-trained vision-language-action models while utilizing significantly smaller architectures.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes