CVAug 2

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

arXiv:2608.013929.8
Predicted impact top 36% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in text-conditioned motion generation, this provides a method to achieve finer alignment without manual annotation, improving grounding and potentially generation quality.

The paper addresses the problem of fine-grained motion-language alignment in text-conditioned human motion generation, where only clip-level supervision is available. They propose FineMoLA, a weakly supervised framework using optimal transport to infer frame-phrase alignments, and show improved motion-text grounding on SnapMoGen.

Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form descriptions typically provide supervision only at the clip level, without explicit temporal correspondence between motion frames and language. This limits fine-grained motion--text grounding and temporally precise generation. We propose FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations. Our method first segments long-form descriptions into action-bearing phrases, and then formulates motion--language alignment as an optimal transport problem, which naturally models many-to-many relations between motion frames and text under global constraints. With entropic regularization and Sinkhorn iterations, FineMoLA efficiently infers pseudo frame-level alignments without human labeling. Experiments on SnapMoGen demonstrate that the learned alignments outperform baselines in motion--text grounding.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes