CVJun 21, 2020

Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020

arXiv:2006.11693v212 citations
AI Analysis

This is an incremental improvement for video understanding tasks, specifically targeting the ActivityNet Challenge 2020.

The paper tackled dense video captioning by proposing a two-stage pipeline for extracting temporal event proposals and generating captions, achieving a 9.28 METEOR score on the test set.

This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: first, we extract a set of temporal event proposals; then we propose a multi-event captioning model to capture the event-level temporal relationships and effectively fuse the multi-modal information. Our approach achieves a 9.28 METEOR score on the test set.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes