SDAug 4

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

arXiv:2608.030219.5h-index: 2
Predicted impact top 42% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers in audio generation, this work addresses the underexplored integration of acoustic priors in neural codecs, offering a method to enhance synthesis fidelity for singing voices.

MeloCodec introduces a neural audio codec that integrates melodic priors via a Tokenize-then-Fuse paradigm to improve singing voice representation. It outperforms baselines in pitch consistency and enables controllable pitch manipulation with minimal timbre degradation.

Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes