Kyunghyun Cho

2papers

2 Papers

1.4CVFeb 1, 2021
Self-Supervised Equivariant Scene Synthesis from Video

Cinjon Resnick, Or Litany, Cosmas Heiß et al.

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into background, characters, and their animations. Our method capitalizes on moving characters being equivariant with respect to their transformation across frames and the background being constant with respect to that same transformation. After training, we can manipulate image encodings in real time to create unseen combinations of the delineated components. As far as we know, we are the first method to perform unsupervised extraction and synthesis of interpretable background, character, and animation. We demonstrate results on three datasets: Moving MNIST with backgrounds, 2D video game sprites, and Fashion Modeling.

1.2CVNov 11, 2020
Learned Equivariant Rendering without Transformation Supervision

Cinjon Resnick, Or Litany, Hugo Larochelle et al.

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into objects and background. Our method relies on moving objects being equivariant with respect to their transformation across frames and the background being constant. After training, we can manipulate and render the scenes in real time to create unseen combinations of objects, transformations, and backgrounds. We show results on moving MNIST with backgrounds.