CVJul 16

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

arXiv:2607.1468126.9h-index: 13Has Code
Predicted impact top 2% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of coordinating multiple visual references in video editing, which is critical for content creators and video editing applications.

ReBind introduces a framework for multi-reference video editing that uses structured instructions with explicit reference tokens to precisely bind visual attributes to their sources, achieving state-of-the-art performance among open-source methods.

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordinate information from multiple visual sources accurately. We identify a critical deficiency in existing approaches. Existing editing instructions lack explicit reference relationships, and most multimodal large language models (MLLMs) cannot generate them reliably. To address this problem, we propose ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing. Our key insight is embedding reference tokens at semantic positions to eliminate ambiguity and establish precise bindings between visual attributes and their sources. We develop ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attributes and their reference sources through a two-stage progressive scheme for precise reference relationships. We further develop ReBind-Edit, which enables lightweight adaptation of text-to-video models to coordinate multiple references by binding visual attributes to their designated sources. Extensive experiments demonstrate that ReBind substantially outperforms general-purpose MLLMs in instruction quality and achieves state-of-the-art performance among open-source methods on reference image conditioned video editing. Our project webpage: https://rebind-mrv2v.github.io/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes