CVMay 6, 2024

MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization

arXiv:2405.03803v15 citations
Originality Incremental advance
AI Analysis

This work addresses the challenge of controlling diversity in human motion generation for applications in animation and robotics, representing an incremental improvement by adapting existing DPO methods with AI feedback.

The paper tackled the problem of aligning text-to-motion diffusion models to generate realistic and text-aligned outputs by proposing MoDiPO, which uses AI feedback for Direct Preference Optimization, resulting in significantly more realistic motions with improved Frechet Inception Distance while maintaining other performance metrics.

Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language conditioning. Their inherent stochasticity, that is the ability to generate various outputs from a single input, is key to their success. However, this diversity should not be unrestricted, as it may lead to unlikely generations. Instead, it should be confined within the boundaries of text-aligned and realistic generations. To address this issue, we propose MoDiPO (Motion Diffusion DPO), a novel methodology that leverages Direct Preference Optimization (DPO) to align text-to-motion models. We streamline the laborious and expensive process of gathering human preferences needed in DPO by leveraging AI feedback instead. This enables us to experiment with novel DPO strategies, using both online and offline generated motion-preference pairs. To foster future research we contribute with a motion-preference dataset which we dub Pick-a-Move. We demonstrate, both qualitatively and quantitatively, that our proposed method yields significantly more realistic motions. In particular, MoDiPO substantially improves Frechet Inception Distance (FID) while retaining the same RPrecision and Multi-Modality performances.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes