CVMMMay 20, 2023

Movie101: A New Movie Understanding Benchmark

arXiv:2305.12140v2232 citationsHas Code
Originality Incremental advance
AI Analysis

This work addresses the need for more realistic benchmarks in movie understanding for visually impaired users, though it is incremental as it builds on existing video captioning tasks.

The authors tackled the problem of automatic movie narration for the visually impaired by constructing a large-scale Chinese movie benchmark, Movie101, which includes a movie clip narrating task and a new evaluation metric, MNScore, that correlates best with human judgment.

To help the visually impaired enjoy movies, automatic movie narrating systems are expected to narrate accurate, coherent, and role-aware plots when there are no speaking lines of actors. Existing works benchmark this challenge as a normal video captioning task via some simplifications, such as removing role names and evaluating narrations with ngram-based metrics, which makes it difficult for automatic systems to meet the needs of real application scenarios. To narrow this gap, we construct a large-scale Chinese movie benchmark, named Movie101. Closer to real scenarios, the Movie Clip Narrating (MCN) task in our benchmark asks models to generate role-aware narration paragraphs for complete movie clips where no actors are speaking. External knowledge, such as role information and movie genres, is also provided for better movie understanding. Besides, we propose a new metric called Movie Narration Score (MNScore) for movie narrating evaluation, which achieves the best correlation with human evaluation. Our benchmark also supports the Temporal Narration Grounding (TNG) task to investigate clip localization given text descriptions. For both two tasks, our proposed methods well leverage external knowledge and outperform carefully designed baselines. The dataset and codes are released at https://github.com/yuezih/Movie101.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes