CVJun 4, 2020

Visually Guided Sound Source Separation using Cascaded Opponent Filter Network

arXiv:2006.03028v212.424 citations

Originality Incremental advance

AI Analysis

This work addresses the problem of separating mixed audio signals using visual information, which is incremental as it builds on existing methods with novel components like opponent filters and sound source location masking.

The paper tackles visually guided sound source separation by proposing a Cascaded Opponent Filter (COF) framework that recursively refines separation using visual cues, achieving state-of-the-art performance on three challenging datasets (MUSIC, A-MUSIC, and A-NATURAL).

The objective of this paper is to recover the original component signals from a mixture audio with the aid of visual cues of the sound sources. Such task is usually referred as visually guided sound source separation. The proposed Cascaded Opponent Filter (COF) framework consists of multiple stages, which recursively refine the source separation. A key element in COF is a novel opponent filter module that identifies and relocates residual components between sources. The system is guided by the appearance and motion of the source, and, for this purpose, we study different representations based on video frames, optical flows, dynamic images, and their combinations. Finally, we propose a Sound Source Location Masking (SSLM) technique, which, together with COF, produces a pixel level mask of the source location. The entire system is trained end-to-end using a large set of unlabelled videos. We compare COF with recent baselines and obtain the state-of-the-art performance in three challenging datasets (MUSIC, A-MUSIC, and A-NATURAL). Project page: https://ly-zhu.github.io/cof-net.

View on arXiv PDF

Similar