CV AIOct 10, 2022

Graph2Vid: Flow graph to Video Grounding for Weakly-supervised Multi-Step Localization

Nikita Dvornik, Isma Hadji, Hai Pham, Dhaivat Bhatt, Brais Martinez, Afsaneh Fazly, Allan D. Jepson

arXiv:2210.04996v26.56 citationsh-index: 11

Originality Incremental advance

AI Analysis

This addresses the problem of reducing annotation requirements for video localization in instructional contexts, though it is incremental as it builds on existing weakly-supervised methods.

The paper tackles weakly-supervised multi-step localization in instructional videos by using generic procedural text to create flow graphs that capture partial order of steps, eliminating the need for step order annotations; experiments on the extended CrossTask dataset show Graph2Vid is more efficient and yields strong localization results without order annotation.

In this work, we consider the problem of weakly-supervised multi-step localization in instructional videos. An established approach to this problem is to rely on a given list of steps. However, in reality, there is often more than one way to execute a procedure successfully, by following the set of steps in slightly varying orders. Thus, for successful localization in a given video, recent works require the actual order of procedure steps in the video, to be provided by human annotators at both training and test times. Instead, here, we only rely on generic procedural text that is not tied to a specific video. We represent the various ways to complete the procedure by transforming the list of instructions into a procedure flow graph which captures the partial order of steps. Using the flow graphs reduces both training and test time annotation requirements. To this end, we introduce the new problem of flow graph to video grounding. In this setup, we seek the optimal step ordering consistent with the procedure flow graph and a given video. To solve this problem, we propose a new algorithm - Graph2Vid - that infers the actual ordering of steps in the video and simultaneously localizes them. To show the advantage of our proposed formulation, we extend the CrossTask dataset with procedure flow graph information. Our experiments show that Graph2Vid is both more efficient than the baselines and yields strong step localization results, without the need for step order annotation.

View on arXiv PDF

Similar