CVJul 1

Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold

arXiv:2607.0064710.6Has Code
Predicted impact top 36% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners using training-free guidance in diffusion models, this work identifies x-prediction as a superior foundation to avoid manifold drift, though the insight is incremental.

The paper shows that x-prediction models keep training-free diffusion guidance on the data manifold more reliably than ε- and v-prediction models, especially at high noise levels, as confirmed by experiments on bird generation and style transfer.

Training-free guidance (TFG) steers a pretrained diffusion model toward a desired attribute at inference. To be effective, this guidance must be applied from the earliest, high-noise steps of sampling. Because its objective (a classifier or energy) is defined on clean images, $ε$- and $v$-prediction models must first estimate the clean image $\hat{x}$ from the noisy state at each step, and the accuracy of that estimate determines how easily guidance drifts off the data manifold. $x$-prediction, a recent alternative, outputs the clean image directly, removing this source of error even at high noise. This is our motivation. We provide a theoretical analysis of how each prediction target shapes this accuracy, and introduce guided-class FID (Child FID), a metric that exposes the manifold damage standard evaluation misses. Experiments on a new fine-grained bird benchmark and on style transfer confirm that $x$-prediction keeps guided samples on the manifold most reliably, making it the strongest foundation for training-free guidance. Code is available at https://github.com/ManLuML/on-manifold-tfg

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes