A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks
For researchers and educators using automatically generated scientific videos, this work highlights a critical gap in evaluation—focusing on explanatory quality rather than content presence—but is incremental as it proposes a new metric rather than a new generation method.
The paper introduces EffectivePresentationScorer, a framework to evaluate whether paper-to-video talks actually teach viewers, and finds that current generation systems fail to explain prerequisite concepts or clarify why methods work, a gap missed by existing metrics.
Automatically generated videos from scientific papers are increasingly used for education and research dissemination. However, existing evaluation metrics mainly measure visual quality or whether key points from the paper appear in the video without assessing whether the video actually helps viewers understand the ideas. We introduce EffectivePresentationScorer, a framework for evaluating the instructional quality of scientific presentation videos. It checks whether a video explains the main ideas clearly, introduces needed background concepts, and connects technical details to the main contribution of the paper. When we apply EffectivePresentationScorer to the existing paper-to-video generation systems, we find that generated videos mention the correct topics and follow the structure of the paper but fail to explain prerequisite concepts or clarify why the method works. These failures are often ignored by existing video evaluation metrics, which focus on content presence rather than explanatory quality.