CLJun 10

Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures

arXiv:2606.12576v115.1
Predicted impact top 66% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for automated, paper-grounded video explanations of complex scientific figures, which is a novel but niche problem for researchers and educators.

The authors tackle the problem of generating narrated, region-grounded walkthrough videos from scientific figures and their papers. Their proposed pipeline, MINARD, outperforms existing methods in both automatic and human evaluations on the new FigTalk benchmark.

Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems and benchmarks. To address this, we introduce paper-grounded figure-to-video generation: generating narrated, region-grounded walkthrough videos from a figure and its paper. We propose MINARD (Multimodal Interpretation of Narrated Architecture via Region Decomposition), a pipeline that generates paper-grounded narrations and sequentially grounds them to figure regions. We also release FigTalk, a benchmark with new sequential and component-level grounding metrics derived. On FigTalk, MINARD generates humanlike, paper-faithful narrations and outperforms narration-conditioned figure spatial grounding compared to existing approaches in both automatic and human evaluation

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes