DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent
For practitioners using pre-trained diffusion models, DiffRGD offers a plug-and-play method to improve controlled generation without retraining.
DiffRGD introduces a distribution-aware guidance framework that preserves the latent Gaussian structure during inference-time diffusion guidance, outperforming previous methods in most image restoration and conditional generation tasks.
Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To control the generation without costly re-training or fine-tuning, many works seek inference-time guidance methods to steer the latent via a differentiable objective at inference time. However, these methods cannot effectively preserve the original Gaussian distribution because they introduce distributional drift, thereby degrading the sample quality. To address this gap, we propose DiffRGD, a distribution-aware guidance framework that explicitly preserves the latent Gaussian structure. DiffRGD formulates each sampling step as a constrained optimization problem on a spherical manifold induced by the latent Gaussian distribution, and solves it efficiently via Riemannian Gradient Descent (RGD). DiffRGD is a plug-and-play method that can be seamlessly integrated into any pre-trained diffusion model. Extensive experiments demonstrate that DiffRGD outperforms previous methods in most image restoration and conditional generation tasks. Our codebase is available at https://github.com/jwliao1209/DiffRGD.