Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration
For researchers and users of diffusion models, this framework provides a novel way to interpret semantic structure evolution, though it is an incremental tool for a known bottleneck.
The paper presents a visual analytics framework to explore attention dynamics in diffusion models, enabling structured analysis of token-level cross-attention maps across generation steps. Case studies on a 60-prompt benchmark show recurring interpretable patterns, facilitating human-AI collaboration.
Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structure remain difficult to interpret. Many existing workflows rely on aggregated attention or scalar summaries that separate temporal change from image-space evidence. To address this gap, we present a visual analytics framework for exploring attention dynamics in diffusion models: the step-indexed evolution of token-level cross-attention maps, their temporal concentration, and their spatial relationships. Our approach enables structured analysis of attention behavior across generation steps by integrating quantitative measures with data-driven stage identification in an interactive workflow. Case studies on a structured 60-prompt Stable-Diffusion-class benchmark illustrate recurring, interpretable patterns within this setting and show how linked temporal and spatial views facilitate the observation and discussion of generative processes, supporting more effective human-AI collaboration.