CVJun 24

H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

arXiv:2606.2557810.2
Predicted impact top 48% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For virtual try-on applications, H-Adapter addresses the challenge of large head-pose discrepancies in hairstyle transfer, offering a more robust solution.

H-Adapter improves pose robustness in hairstyle transfer by using a region-specific loss to derive source-aligned hair masks for diffusion-based inpainting, achieving best FID, FID_CLIP, and CLIP-I under pose differences while maintaining competitive non-hair preservation.

Hairstyle transfer has practical applications such as virtual try-on, yet remains challenging when the source and reference exhibit large head-pose discrepancies. We propose H-Adapter, which improves pose robustness by training with a region-specific loss that disentangles hair and non-hair objectives and thereby induces spatially disentangled cross-attention, from which a source-aligned hair edit mask is derived to guide diffusion-based inpainting. Experiments on pose-agnostic and pose-different subsets demonstrate strong quantitative results, including the best FID, $\mathrm{FID}_{\mathrm{CLIP}}$, and CLIP-I under pose differences, while maintaining competitive non-hair preservation and improving qualitative fidelity to fine-grained reference hairstyle details. Beyond source-conditioned transfer, H-Adapter supports practical extensions including text-to-image generation, auxiliary prompt-based hair color control, and compatibility with an identity-preserving IP-Adapter variant. We also introduce a VLM-as-a-judge protocol and observe consistent gains in hairstyle faithfulness, non-hair preservation, and artifact quality.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes