CVAILGIVOct 17, 2024

Movie Gen: A Cast of Media Foundation Models

Meta AI
arXiv:2410.13720v2554 citationsh-index: 30
Originality Highly original
AI Analysis

This work addresses the challenge of comprehensive media generation for applications in entertainment, education, and communication, representing a significant advancement rather than an incremental improvement.

The authors tackled the problem of generating high-quality 1080p HD videos with synchronized audio and additional capabilities like instruction-based editing and personalization, achieving state-of-the-art results on multiple tasks including text-to-video synthesis and video-to-audio generation with a 30B parameter model producing 16-second videos.

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and generation of personalized videos based on a user's image. Our models set a new state-of-the-art on multiple tasks: text-to-video synthesis, video personalization, video editing, video-to-audio generation, and text-to-audio generation. Our largest video generation model is a 30B parameter transformer trained with a maximum context length of 73K video tokens, corresponding to a generated video of 16 seconds at 16 frames-per-second. We show multiple technical innovations and simplifications on the architecture, latent spaces, training objectives and recipes, data curation, evaluation protocols, parallelization techniques, and inference optimizations that allow us to reap the benefits of scaling pre-training data, model size, and training compute for training large scale media generation models. We hope this paper helps the research community to accelerate progress and innovation in media generation models. All videos from this paper are available at https://go.fb.me/MovieGenResearchVideos.

Code Implementations2 repos
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes