CVJun 21, 2017

Learnable pooling with Context Gating for video classification

arXiv:1706.06905v231.6344 citationsHas Code

Originality Incremental advance

AI Analysis

This work addresses video classification for large-scale multi-modal datasets, representing an incremental improvement over existing temporal aggregation methods.

The paper tackles video classification by proposing a learnable non-linear unit called Context Gating to model interdependencies in network activations, along with clustering-based aggregation and a two-stream architecture for audio-visual features, achieving state-of-the-art results on the Youtube-8M v2 dataset.

Current methods for video analysis often extract frame-level features using pre-trained convolutional neural networks (CNNs). Such features are then aggregated over time e.g., by simple temporal averaging or more sophisticated recurrent neural networks such as long short-term memory (LSTM) or gated recurrent units (GRU). In this work we revise existing video representations and study alternative methods for temporal aggregation. We first explore clustering-based aggregation layers and propose a two-stream architecture aggregating audio and visual features. We then introduce a learnable non-linear unit, named Context Gating, aiming to model interdependencies among network activations. Our experimental results show the advantage of both improvements for the task of video classification. In particular, we evaluate our method on the large-scale multi-modal Youtube-8M v2 dataset and outperform all other methods in the Youtube 8M Large-Scale Video Understanding challenge.

View on arXiv PDF Code

Similar