CVJul 1

Training-Free Debiasing of Diffusion Models via CLIP-Guided Denoising Optimization

arXiv:2607.008179.2
Predicted impact top 44% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners using diffusion models, TES offers a practical, scalable solution to mitigate demographic bias without costly retraining or quality degradation.

The paper tackles demographic bias in text-to-image diffusion models, proposing a training-free framework (TES) that optimizes text embeddings during inference to reduce bias. TES outperforms existing training-free methods in fairness while maintaining competitive image quality.

Text-to-image diffusion models achieve impressive visual quality, yet demographic bias remains a challenge, as neutral prompts consistently produce stereotypical representations across gender and race. Existing approaches remain limited by costly retraining or by inference-time interventions that often degrade image quality and semantic alignment. We propose Text Embedding Steering (TES), a training-free framework that mitigates demographic bias by directly optimizing conditional text embeddings during the diffusion process. We show that a two-stage strategy - early-stage global alignment followed by iterative denoising-time refinement with CLIP-based feedback - enables stable and controllable attribute steering without modifying model parameters. Extensive experiments on Stable Diffusion demonstrate that TES outperforms existing training-free baselines in fairness while maintaining competitive image quality. These results highlight that inference-time text embedding optimization is a practical and scalable solution for fairness-aware generation in diffusion models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes