LGAIJul 10

From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime

arXiv:2607.2054012.8
Predicted impact top 15% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a theoretical foundation for noise schedule design in diffusion models, addressing a key practical bottleneck for practitioners.

The authors develop a statistical framework for optimal noise-level allocation in diffusion training, showing that the optimized schedule is atomic (concentrated on finitely many noise levels) under convexity assumptions, and propose a square-root entropy scheduling proxy that improves training efficiency on discrete domains and matches heuristics on continuous images.

How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise schedules are based largely on heuristics or empirical tuning. Here, we develop a general statistical framework for studying asymptotically optimal noise-level allocation in diffusion training. Our first main result concerns the fully coupled regime, where information can spread between different time points. Under convexity or Polyak-Lojasiewicz-type assumptions, we show that the optimized training schedule admits an atomic minimizer, concentrated on finitely many noise levels. Our second main result specializes this framework to an idealized independent-learner regime, intended to model temporal specialization in neural networks. Under an additional feature-noise decoupling condition, a random-matrix analysis leads to an information-theoretic proxy: the decoupled sampling density is proportional to the square root of the generative entropy rate, the rate at which conditional entropy grows along the forward process. We test these predictions in controlled settings where the coupled objective can be optimized directly, including Dirac mixtures, low-dimensional manifolds, and MNIST. In these settings, the optimized schedules are consistently finite-support, while the smooth entropic proxy closely tracks the atomic optimum in neural-network models and breaks down mainly in the fully coupled parametric case, as the theory suggests. We then evaluate the entropic schedule in larger-scale experiments, where full schedule optimization is currently intractable. The results indicate that square-root entropy scheduling can substantially improve training efficiency on discrete domains and remains competitive with standard EDM-style heuristics on continuous images.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes