ASCVJul 16

WanSong v1.0 Technical Report

arXiv:2607.1474915.9h-index: 10
Predicted impact top 21% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a simpler, more efficient alternative to autoregressive and cascaded pipelines for commercial-grade song generation, enabling faster inference and easier customization.

WanSong is a pure diffusion-based model for generating high-fidelity, multilingual songs up to 5 minutes with dual stems in a single run, achieving faster inference through step-distillation. It addresses challenges in efficient generation and controllability for long-form audio.

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes