CVJun 13

RefGC-SR$^2$: Reference-guided Generated Content Super-Resolution and Refinement

arXiv:2606.1515817.3
Predicted impact top 19% in CV · last 90 daysOriginality Highly original
AI Analysis

For users of reference-guided generation pipelines, this work addresses the critical bottleneck of losing fine-grained details from high-resolution references and introduces artifacts, enabling higher-quality and more usable outputs.

The paper introduces a new task, RefGC-SR², which simultaneously recovers lost details, refines generative artifacts, and upscales output in reference-guided generation. The proposed frequency-aware diffusion transformer model significantly outperforms existing RefGCR and RefSR baselines in both identity fidelity and resolution recovery.

Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation: the object-centric high-resolution reference image (HRRI) provided by users is downsampled to a fixed low-resolution (LR) before being fed into the model, so the fine-grained details are discarded before the output is even produced. In addition, the generation step then introduces its own artifacts (e.g., identity distortion) on top of this loss. Existing reference-guided generated content refinement (RefGCR) methods can correct some of these artifacts but still operate in the LR domain; reference-guided super-resolution (RefSR) methods recover resolution but assume natural-image degradations and ignore the artifact distribution of generative pipelines. To address both gaps in a single formulation, we introduce a new task: reference-guided generated content super-resolution-refinement (RefGC-SR$^2$), where the original HRRI is reused at the post-processing stage to recover lost details, refine generative artifacts, and upscale the output simultaneously. We construct the first real-world triplet data generation pipeline for this RefGC-SR$^2$ task, training a diptych-conditioned generator to synthesize paired low-quality anchors that public pretrained models cannot provide. We further present a frequency-aware diffusion transformer model for RefGC-SR$^2$ that selectively injects fine details from the HRRI while removing generative artifacts. Extensive experiments demonstrate that our RefGC-SR$^2$ model successfully (i) refines the object identity faithfully with respect to the reference, and (ii) recovers high-resolution details, so that the final result is significantly higher quality and practically more usable compared to existing RefGCR and RefSR baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes